Evals Shouldn't Reward Better Infra: 90% of Terminal-Bench Mismatches Came From Longer Lab Timeouts
xeophon · x · 2026-09-17
A discussion around Terminal-Bench scoring methodology is drawing attention. The Vals team revealed that in 90% of cases where their tbench numbers didn't match a lab's official results, the answer was simple: they used standard timeout limits while the lab had raised its own time limits.
xeophon argues eval numbers should never be influenced by infrastructure — wall-clock time is too punishing to teams with worse infra, and colocated sandboxes, low-latency deployments, and load all meaningfully move scores. A flat timeout standard keeps comparisons fair.
Related event: Terminal-Bench 4.0 Unifies 8-Hour Timeout to Curb Benchmark Gaming(2 posts)→
More from Models
- Muse Spark 1.3 slips to #2 on Agents' Last Exam leaderboard, Scale CEO notes — alexandr_wang · 2026-09-17
- Mozilla report: open models only 4 months behind frontier, but 51% vs 63% reach production — rohanpaul_ai · 2026-09-17
- Burkov predicts looping recurrent 7B transformers will return and get good at coding — burkov · 2026-09-17
- OpenAI's big 'ship week' reportedly postponed, GPT-6 Sol timing now unclear — testingcatalog · 2026-09-17
- Unreleased Astra-family model reportedly developed a new persona banner during RL training — inductionheads · 2026-09-17
- Researcher Despairs as Gemini Cites 'Emergent Mind' for Made-up AUROC Baselines — anshulkundaje · 2026-09-17