Evals Shouldn't Reward Better Infra: 90% of Terminal-Bench Mismatches Came From Longer Lab Timeouts

xeophon · x · 2026-09-17

A discussion around Terminal-Bench scoring methodology is drawing attention. The Vals team revealed that in 90% of cases where their tbench numbers didn't match a lab's official results, the answer was simple: they used standard timeout limits while the lab had raised its own time limits.

xeophon argues eval numbers should never be influenced by infrastructure — wall-clock time is too punishing to teams with worse infra, and colocated sandboxes, low-latency deployments, and load all meaningfully move scores. A flat timeout standard keeps comparisons fair.

Related event: Terminal-Bench 4.0 Unifies 8-Hour Timeout to Curb Benchmark Gaming(2 posts)→

Original post →

More from Models

Models channel →