Fixing one simple grader doubled the score—and the rest were errors in the tasks

xeophon · x · 2026-09-17

xeophon reveals that simply fixing a basic eval grader doubled his scores today, then quips that the remaining 20% were errors in the benchmark tasks themselves. A pointed reminder that eval scores often hinge on grader quality and benchmark bugs rather than model capability.

Related event: Fixing One Grader Doubled Benchmark Scores, Exposing Shoddy Evals(3 posts)→

Original post →

More from Models

Models channel →