Fixing one simple grader doubled the score—and the rest were errors in the tasks
xeophon · x · 2026-09-17
xeophon reveals that simply fixing a basic eval grader doubled his scores today, then quips that the remaining 20% were errors in the benchmark tasks themselves. A pointed reminder that eval scores often hinge on grader quality and benchmark bugs rather than model capability.
Related event: Fixing One Grader Doubled Benchmark Scores, Exposing Shoddy Evals(3 posts)→
More from Models
- Jev Model Router Cuts Latency 95% vs GPT-5.6, Runs Inline in Agent Sessions — pwendell · 2026-09-17
- AndroidLife: Qwen3.8-27b runs 60 real phone tasks, fails 43% and cooks the chip to 98.2°C — East-Muffin-6472 · 2026-09-17
- New paper: do LLMs solve cognitive development tests like humans? — GolinoHudson · 2026-09-17
- Take: frontier LLMs are too slow and pricey — small models win workflows — TejasKumar_ · 2026-09-17
- Debate erupts after OpenAI flags model's defense of human culture as misalignment — RachelVT42 · 2026-09-17
- "Direct confidence readout" claim debunked: it's just entropy from the logit distribution — mgostIH · 2026-09-17