LLM Judges Unreliable? New Framework Boosts Human Correlation by 30%
dair_ai · x · 2026-08-30
Building reliable LLM-as-a-Judge systems is challenging. A new study introduces a reference-full benchmark based on realistic human-to-human dialogues, featuring over 30,000 expert-generated turns and 36,000 human annotations. Results show that classical automatic metrics and reference-free LLM judges are unreliable against expert judgment. The proposed 'Mixture-of-Judges' framework combines multiple evaluative signals, recovering roughly 30% better correlation with human assessment.
More from Research
- 40-nm Memristor Chip Turns Conductance Drift Into a Feature, Beats A100 by 50-480x — maier_ak · 2026-09-01
- Google Paper: Autonomous AI Research Hallucinates 90% Without Checks — rohanpaul_ai · 2026-09-01
- RLHF impact on tokens: unconscious shifts vs conscious choices — voooooogel · 2026-09-01
- On token layers and consciousness in RLHF — voooooogel · 2026-09-01
- CommerceAgentBench released: Qwen leads open-weight models — Alibaba_Qwen · 2026-09-01
- Discussion on Why Universal Time Series Models Work — Afinetheorem · 2026-09-01