Fixing One Grader Doubled Benchmark Scores, Exposing Shoddy Evals
Xeophon reports that fixing a single poorly written grader doubled a model's benchmark score, with the remaining 20% gap caused by task errors — fueling criticism that LLM evals are routinely flawed and unreliable.
2026-09-17 ~ 2026-09-17 · 3 related posts
- Another "hard" benchmark falls — Xeophon mocks recurring eval hype cycles — xeophon · 2026-09-17
- Doubled Benchmark Scores by Fixing One Simple Grader — xeophon · 2026-09-17
- Fixing one simple grader doubled the score—and the rest were errors in the tasks — xeophon · 2026-09-17