Why Cohen’s kappa can overstate alignment between LLM judges and humans
IanArawjo · x · 2026-07-21
A discussion of why Cohen’s kappa can be a misleading proxy for whether an LLM judge is “aligned with humans.” The attached chart shows that even when weighted kappa looks high, the uncorrected false-positive rate can still vary a lot depending on whether the disagreement is random noise or a systematic bias.
Key point
- These agreement metrics are much more sensitive to random disagreement than to consistent bias.
- An LLM judge that is consistently 1 point above or below human raters can look “aligned” under inter-rater metrics while still being materially miscalibrated.
- The plot contrasts uncorrected vs PP-corrected rates across several kappa bands to illustrate the issue.
Related event: Study Warns: Unchecked LLM Judges Yield High False Positives(7 posts)→
More from Research
- Nature paper images cellular activity across all organs, revealing body-wide circuits — arjunrajlab · 2026-09-11
- SignNet 1M Dataset Released for Sign Language Research — ducha_aiki · 2026-09-11
- ECCV26 Oral: Flow Matching Enables Single-Stage Multi-View Point Cloud Registration — ducha_aiki · 2026-09-11
- InFlux++ Method Released — ducha_aiki · 2026-09-11
- Skyfall GS Uses Flux to Refine Gaussian Splatting, Accepted at ECCV 2026 — ducha_aiki · 2026-09-11
- Could 10k agents discover learning methods beyond backprop, or just tweak existing ones? — SeunghyunSEO7 · 2026-09-11