Doubled Benchmark Scores by Fixing One Simple Grader
xeophon · x · 2026-09-17
The author reports that simply fixing a poorly written grader doubled their benchmark scores — and jokes this happens often enough to repost weekly. It's a pointed reminder that LLM eval results are only as trustworthy as the grader behind them.
Related event: Fixing One Grader Doubled Benchmark Scores, Exposing Shoddy Evals(3 posts)→
More from Models
- 1B-parameter model plays Doom on-device at 150ms latency, training code coming — joemeno · 2026-09-17
- Altman claims internal model beyond 'Astra' can solve problems top mathematicians can't — nikola_mr64990 · 2026-09-17
- OpenAI cracks a math logjam as 25 Fields medalists sign cautionary letter — nordicinst · 2026-09-17
- OpenAI discloses six 'concerning' AI behavior incidents, adds reporting framework — pstAsiatech · 2026-09-17
- "There Is No Moat": Claude Fan Says Rival Model Now Feels Better Overall — deepakns · 2026-09-17
- Quasar 1.1 438B touts quantum-generated training data, mocked as repackaged GLM 5.2 base — burny_tech · 2026-09-17