Why a high kappa score still does not make an LLM judge trustworthy

IanArawjo · x · 2026-07-21

A follow-up explanation of the same metric issue: agreement statistics are more sensitive to **random disagreements** than to a **consistent offset**. If an LLM is biased by a nearly constant amount — for example, always scoring one point above or below human raters — inter-rater alignment metrics may still look fine. That means a high kappa score does not necessarily mean the judge is truly reliable. ### Takeaway - Random noise hurts agreement scores more than stable bias. - A judge can look “aligned” while still being systematically off in practice.

Related event: Study Warns of High False Positive Rates in Unchecked LLM Judges(7 posts)→

Original post →

More from Research

Research channel →