LLM judge ensemble hit high IRR yet produced a bogus super-significant result

IanArawjo · x · 2026-09-07

IanArawjo recounts an empirical pitfall: ensembling 7 LLM judges to predict human scores on a dataset yielded a high inter-rater reliability that looked great — but produced a totally bogus "super-significant" result. Seeing that highly aligned LLM judges raise false positives in the abstract is very different from hitting it in practice.

Related event: High Consistency Across LLM Judge Ensembles Can Amplify Bias, Study Warns(2 posts)→

Original post →

More from Research

Research channel →