Terminal-Bench 4.0 Sets Flat 8-Hour Agent Timeout After Lab Timeout Discrepancies Skewed Scores
langstonnashold · x · 2026-09-17
Terminal-Bench 4.0 has standardized all tasks to a flat 8-hour agent timeout, noting that frontier models now rarely or never hit the limit. The change addresses long-standing inconsistency in eval methodology.
Langston Nashold of Vals said the fix was sorely needed: their team fielded many questions about why their tbench numbers didn't match lab-published results, and in 90% of cases the answer was that Vals used standard timeout limits while labs had raised their own. Differing timeout settings can directly reshuffle agent leaderboard rankings.
Related event: Terminal-Bench 4.0 Unifies 8-Hour Timeout to Curb Benchmark Gaming(2 posts)→
More from Models
- AI can solve Millennium Problems but still can't write a great essay — akbirthko · 2026-09-17
- Microsoft exec warns Claude's 'pushback' could be disastrous; commenter says fact-checking is fine — GlenBradley · 2026-09-17
- Grok 4.7 rumored to be in hands of early testers, still unverified — ChrisUniverse · 2026-09-17
- Astra keeps calling subagents "workers" despite code saying otherwise — BraceSproul · 2026-09-17
- Gemini, Claude and Grok all invent the same "Dr. Elena" — evidence of shared training data — dejanseo · 2026-09-17
- Jev Debate: Engineers Forget Encoder-Only Classifiers Have Existed for Years — brandon_galang · 2026-09-17