LLM-as-Judge Evaluation Reliability and Systematic Biases
Using Cohen's kappa reveals LLM judges are reliably wrong, not reliably right.
Leila Zola
Section
1 story in Benchmark Evaluation.
Using Cohen's kappa reveals LLM judges are reliably wrong, not reliably right.