Terminal-Bench 4.0 scoring flaws hit <3% of rollouts, despite Epoch finding 30/66 defective tasks
DimitrisPapail · x · 2026-09-19
Epoch's review found 30 of 66 Terminal-Bench 4.0 tasks have scoring defects (exploitable graders, answer leakage, correct solutions rejected), sparking "TB 4.0 is broken" takes.
@DimitrisPapail argues Epoch's framing, while not literally saying broken, implies distrust. @ajratner adds key context: the terminal-bench team flagged these issues themselves as PRs (how Epoch found them), and empirically they affect <3% of leaderboard rollouts — well within reported CIs.
Consensus: exploitable graders need fixing, but "30/66 tasks have issues" ≠ "45% of results are wrong" — exploitability doesn't tell you how often it happened. The team is advancing continuously improved open benchmarks rather than shipping a broken one.
Related event: Audit Sparks Debate Over Terminal Bench 4.0 Flaws(3 posts)→
More from Research
- Economist John Horton uses AI voice interviews to verify authors understand their own papers — soumitrashukla9 · 2026-09-19
- Dual-system RL agent crushes WC3 insane computer, open-source soon; 'bring back Dota' — generativist · 2026-09-19
- Imagined songs decoded from intracranial brain signals via relative pitch — DrKavner · 2026-09-19
- AI splits research into three tribes: gatekeepers, AI cultists, and curators — ipeirotis · 2026-09-19
- Simulation of Jupiter's magnetosphere using full magnetohydrodynamics equations — burny_tech · 2026-09-19
- Four LLM families rank the top 500 open math problems in 34,890 pairwise judgments — zero0_one1 · 2026-09-19