Mark Tenenholtz: Unsat Benchmarks Over 3 Months Old Are Just Badly Built
marktenenholtz · x · 2026-10-07
Mark Tenenholtz argues that most benchmarks older than 2-3 months that aren't saturated only have headroom because the tasks are badly written, the grading is poorly implemented, or both — not because models genuinely can't do better.
More from Models
- GitHub Copilot embraces local AI at Microsoft Windows launch, touting llama.cpp — film_girl · 2026-10-08
- Aleph Alpha Open-Sources Kolibri, a 78B-Parameter German-Reasoning Model Under Apache 2.0 — lmoroney · 2026-10-08
- User mocks ChatGPT error after latest update: 'guess I'll go touch grass' — TheMoonMidas · 2026-10-08
- Agent Arena: Jev Router costs 38% more than DeepSeek V4.1 Flash but matches Opus 5.5 steerability — arena · 2026-10-08
- Claude Haiku 5.5 hits Databricks Day 0: ~15% better than Haiku 4.5 at a fraction of the cost — matei_zaharia · 2026-10-08
- Perplexity CEO Shows Off Decider Model Speedrunning 'Who Wants to Be a Millionaire?' — AravSrinivas · 2026-10-08