Coding Benchmarks Are Following Trivia Evals Into Oblivion, Argues Xeophon
xeophon · x · 2026-09-18
Xeophon observes that entire eval categories die as models improve: nobody asks trivia anymore and multiple-choice knowledge benchmarks are dead.
He increasingly feels the same about coding benchmarks — their usefulness as an evaluation signal is fading.
More from Models
- Databricks claims its agents match Claude and GPT-5.6 Luna at twice the speed, per its own tests — emmanuelvivier · 2026-09-18
- Dev surveys Jev demos: self-driving cars, rockets, new languages and more — iannuttall · 2026-09-18
- What should we call the class of models competing with JEV? — altryne · 2026-09-18
- Armin Ronacher floats replacing MCP with codemode + OpenAPI + RAG over API docs — mitsuhiko · 2026-09-18
- Opus 4.8 predicted to win a cult following among devs like GPT-4o did — RileyRalmuto · 2026-09-18
- Chinese coding models use 2-3x the tokens, so cheaper per token isn't cheaper overall — craigbalding · 2026-09-18