Terminal-Bench 4.0 Sets Flat 8-Hour Agent Timeout After Lab Timeout Discrepancies Skewed Scores

langstonnashold · x · 2026-09-17

Terminal-Bench 4.0 has standardized all tasks to a flat 8-hour agent timeout, noting that frontier models now rarely or never hit the limit. The change addresses long-standing inconsistency in eval methodology.

Langston Nashold of Vals said the fix was sorely needed: their team fielded many questions about why their tbench numbers didn't match lab-published results, and in 90% of cases the answer was that Vals used standard timeout limits while labs had raised their own. Differing timeout settings can directly reshuffle agent leaderboard rankings.

Related event: Terminal-Bench 4.0 Unifies 8-Hour Timeout to Curb Benchmark Gaming(2 posts)→

Original post →

More from Models

Models channel →