AI Judges Overestimate New Models and Miss Hidden Flaws
rohanpaul_ai · x · 2026-07-06
Analysis indicates that while using AI as a judge (LLM-as-judge) can directionally track model progress, it often misses hidden quality issues that human evaluators would catch.
Automated judges showed only about a 3% agreement rate with humans on early models but severely overestimate newer ones: the human pass rate for GPT-5.5 was inflated from 6.25% to 17.9%, and Opus 4.8 from 8.33% to 18.8%. This suggests that automated evaluations may suffer from systematic distortion when assessing new models.
More from Models
- Kimi K3 may be strong on cyber, but token efficiency keeps it off UK AISIS — teortaxesTex · 2026-07-27
- European ChatGPT Plus users are now seeing an “Extra High” quality option — PressPlayPlease7 · 2026-07-27
- Opus 5 reportedly aces a car-racing game test on the first try — soumitrashukla9 · 2026-07-27
- Claude Opus 5 arrives at half the price and tops Frontier-Bench claims — GregCook2011 · 2026-07-27
- Open models may beat closed ones for cyber defense, researchers argue as Kimi K3 impresses — eliebakouch · 2026-07-27
- Opus 5 notices when its own generated game looks bad — Angaisb_ · 2026-07-27