FULL STORY

Epoch AI Audits 15 Benchmarks, Only 4 Pass

Epoch AI launched a systematic audit of AI benchmarks, passing only 4 of 15. Terminal Bench 4.0 was found largely corrupted, sparking debate over its impact on model rankings.

2026-09-18 ~ 2026-09-19 · 2 episodes · 20 posts

Episode 1 · Epoch AI's Benchmark Reviews: Only 4 of 15 AI Benchmarks Pass Audit, 9 Found Flawed (2026-09-18, 17 posts)

Epoch AI announced on September 18 the launch of its Benchmark Reviews program, kicking off a systematic audit of AI benchmarks. The first batch covered 15 popular benchmarks, with the conclusion that only 4 earned "Verified" status, 9 were judged flawed, and 2 could not be fully assessed due to insufficient information. The results show that most of the mainstream benchmarks widely cited across the industry failed the audit, posing a direct challenge to the credibility of model capability evaluations.

Confirmed

  • Epoch AI launched the Benchmark Reviews program to systematically audit AI benchmarks
  • The first batch covered 15 benchmarks: 4 Verified, 9 Flawed, and 2 unassessable due to insufficient information
  • The audit found specific issues in some benchmarks — for example, 46% of sampled questions in Humanity's Last Exam (HLE) were found to have problems
  • Terminal Bench 4.0 is also among the benchmarks flagged with issues

Why it matters

  • Benchmark scores underpin model capability evaluations and leaderboard rankings, so most mainstream benchmarks failing the audit casts doubt on the credibility of many current model comparisons
  • The program gives the community a public reference for distinguishing trustworthy from untrustworthy benchmarks, potentially driving improvements in benchmark design
  • The disclosure that widely used benchmarks like HLE contain a high proportion of problematic questions is a reminder to pay attention to data quality when citing scores

Episode 2 · Audit Sparks Debate Over Terminal Bench 4.0 Flaws (2026-09-19, 3 posts)

An audit flagged 45.5% of Terminal Bench 4.0 tasks as broken, and Epoch found scoring flaws in 30 of 66 tasks, but maintainer AlexGDimakis countered that known issues affect less than 3% of the leaderboard, fueling debate over the benchmark's integrity.