Spearman and Kappa metrics misleading for LLM judge QC

IanArawjo · x · 2026-08-25

Ian Arawjo extends his critique to common metrics like Spearman's rho, Krippendorf's alpha, and Kendall's tau, recommended by guidelines such as those used by OpenAI. He warns that without bias correction, these metrics are totally misleading as quality control gates for LLM-as-a-Judge systems.

Related event: High Inter-Rater Consistency in LLM Evaluation May Be Misleading, Researcher Warns(2 posts)→

Original post →

More from Research

Research channel →