LLM-as-Judge Evaluation Reliability and Systematic Biases
Using Cohen's kappa reveals LLM judges are reliably wrong, not reliably right.
Leila Zola
Staff Writer
Leila Zola is a staff writer at Expert Data Review covering benchmark evaluation. Based in Melbourne, Leila has written for Expert Data Review since 2022.
1 story · Melbourne
Using Cohen's kappa reveals LLM judges are reliably wrong, not reliably right.