LLM judges only improve statistical power when human alignment reaches about 50%
IanArawjo · x · 2026-07-27
The post says LLM judges only improve statistical power after bias correction when their alignment with humans is roughly 0.50 or higher. Below that threshold, the author argues, researchers should not report LLM-judge results and should instead rely on the smaller set of human labels to estimate significance.
The attached figure shows that the gain depends strongly on judge alignment: around 70% alignment produces clear power gains, while weaker agreement can erase the benefit entirely. The broader claim is that reporting judge scores with low agreement can be misleading, because corrected human+LLM testing only works when the judge is sufficiently aligned with human judgments.
Related event: LLM Judges Require Over 50% Human Alignment for Statistical Efficacy(3 posts)→
More from Research
- Researchers observe models often need an a→b→c→d path before discovering a simple arithmetic trick — dejavucoder · 2026-07-27
- Claude Code is not reliable enough for long research projects without human supervision — _akpiper · 2026-07-27
- Alibaba should ship multiple Qwen sizes, says researcher focused on interpretability — traviscline · 2026-07-27
- AI-assisted vulnerability research for real-time operating systems gets a spotlight — cyb3rops · 2026-07-27
- NVIDIA says AdamW hits a scale ceiling as SOAP and Muon beat it on trillion-token runs — omarsar0 · 2026-07-27
- A proposal calls for LLM prompts to replace peer review’s first pass — ChenhaoTan · 2026-07-27