Epoch AI's benchmark audit calls 9 of 15 flawed; TB4 authors push back

DimitrisPapail · x · 2026-09-19

Epoch AI launched Benchmark Reviews, auditing 15 AI benchmarks: 4 Verified, 9 Flawed, 2 under-specified. After its review found 45.5% of Terminal Bench 4.0 tasks flawed, Berkeley professor Alex Dimakis objected that TB4's bugs are mostly known public GitHub issues being actively fixed, affect under 3% of leaderboard rollouts, and that calling an open-source benchmark "broken" over unfixed issues is unfair—useful audits should surface new issues instead.

Related event: Epoch AI's First Benchmark Audit: 9 of 15 Benchmarks Found Flawed(15 posts)→

Original post →

More from Models

Models channel →