Audit flags 45.5% of Terminal Bench 4.0 tasks as broken; maintainers push back on rigor

DimitrisPapail · x · 2026-09-19

An audit of 15 benchmarks labeled 9 as flawed: 45.5% of Terminal Bench 4.0 tasks were judged broken after reviewing GitHub issues, 46% of 48 randomly sampled HLE questions were broken, and a bug in DeepSWE 1.1 could break grading for every task.

Terminal Bench's Ryan Marten pushed back, arguing the flaws affect <3% of leaderboard rollouts and wouldn't move scores beyond reported confidence intervals. He noted the "flaws" were pulled from already-triaged public GitHub issues — no new issues were filed — and that TB publishes raw receipts (per-trial logs, trajectories, rewards) for every run. His point: all software tasks have bugs; what matters is impact magnitude and the mechanism for catching them.

Related event: Audit Sparks Debate Over Terminal Bench 4.0 Flaws(3 posts)→

Original post →

More from Research

Research channel →