Terminal-Bench 4.0 scoring flaws hit <3% of rollouts, despite Epoch finding 30/66 defective tasks

DimitrisPapail · x · 2026-09-19

Epoch's review found 30 of 66 Terminal-Bench 4.0 tasks have scoring defects (exploitable graders, answer leakage, correct solutions rejected), sparking "TB 4.0 is broken" takes.

@DimitrisPapail argues Epoch's framing, while not literally saying broken, implies distrust. @ajratner adds key context: the terminal-bench team flagged these issues themselves as PRs (how Epoch found them), and empirically they affect <3% of leaderboard rollouts — well within reported CIs.

Consensus: exploitable graders need fixing, but "30/66 tasks have issues" ≠ "45% of results are wrong" — exploitability doesn't tell you how often it happened. The team is advancing continuously improved open benchmarks rather than shipping a broken one.

Related event: Audit Sparks Debate Over Terminal Bench 4.0 Flaws(3 posts)→

Original post →

More from Research

Research channel →