Audit flags 45.5% of Terminal Bench 4.0 tasks as broken; maintainers push back on rigor
DimitrisPapail · x · 2026-09-19
An audit of 15 benchmarks labeled 9 as flawed: 45.5% of Terminal Bench 4.0 tasks were judged broken after reviewing GitHub issues, 46% of 48 randomly sampled HLE questions were broken, and a bug in DeepSWE 1.1 could break grading for every task.
Terminal Bench's Ryan Marten pushed back, arguing the flaws affect <3% of leaderboard rollouts and wouldn't move scores beyond reported confidence intervals. He noted the "flaws" were pulled from already-triaged public GitHub issues — no new issues were filed — and that TB publishes raw receipts (per-trial logs, trajectories, rewards) for every run. His point: all software tasks have bugs; what matters is impact magnitude and the mechanism for catching them.
Related event: Audit Sparks Debate Over Terminal Bench 4.0 Flaws(3 posts)→
More from Research
- AgentZip paper: template-aware memory compression shrinks agent sandbox memory 8.7x — rohanpaul_ai · 2026-09-19
- AgentZip: memory compression for parallel agent sandboxes cuts memory up to 8.7x — rohanpaul_ai · 2026-09-19
- Economist John Horton uses AI voice interviews to verify authors understand their own papers — soumitrashukla9 · 2026-09-19
- Dual-system RL agent crushes WC3 insane computer, open-source soon; 'bring back Dota' — generativist · 2026-09-19
- Imagined songs decoded from intracranial brain signals via relative pitch — DrKavner · 2026-09-19
- AI splits research into three tribes: gatekeepers, AI cultists, and curators — ipeirotis · 2026-09-19