AI Judges Overestimate New Models and Miss Hidden Flaws

rohanpaul_ai · x · 2026-07-06

Analysis indicates that while using AI as a judge (LLM-as-judge) can directionally track model progress, it often misses hidden quality issues that human evaluators would catch.

Automated judges showed only about a 3% agreement rate with humans on early models but severely overestimate newer ones: the human pass rate for GPT-5.5 was inflated from 6.25% to 17.9%, and Opus 4.8 from 8.33% to 18.8%. This suggests that automated evaluations may suffer from systematic distortion when assessing new models.

Original post →

More from Models

Models channel →