Fixing One Grader Doubled Benchmark Scores, Exposing Shoddy Evals

Xeophon reports that fixing a single poorly written grader doubled a model's benchmark score, with the remaining 20% gap caused by task errors — fueling criticism that LLM evals are routinely flawed and unreliable.

2026-09-17 ~ 2026-09-17 · 3 related posts