LLM-as-Judge Evaluation Reliability and Systematic Biases
Using Cohen's kappa reveals LLM judges are reliably wrong, not reliably right.
Marcus Oduya
Staff Writer
Marcus Oduya came to data journalism through a research background in computational linguistics, where he spent years designing evaluation suites for NLP systems. His reporting centers on how benchmarks are constructed, gamed, and reformed across academic and industry settings.
1 story
Using Cohen's kappa reveals LLM judges are reliably wrong, not reliably right.