LLM Judges Unreliable? New Framework Boosts Human Correlation by 30%

dair_ai · x · 2026-08-30

Building reliable LLM-as-a-Judge systems is challenging. A new study introduces a reference-full benchmark based on realistic human-to-human dialogues, featuring over 30,000 expert-generated turns and 36,000 human annotations. Results show that classical automatic metrics and reference-free LLM judges are unreliable against expert judgment. The proposed 'Mixture-of-Judges' framework combines multiple evaluative signals, recovering roughly 30% better correlation with human assessment.

Original post →

More from Research

Research channel →