FULL STORY
Epoch AI Audits 15 Benchmarks, Only 4 Pass
Epoch AI launched a systematic audit of AI benchmarks, passing only 4 of 15. Terminal Bench 4.0 was found largely corrupted, sparking debate over its impact on model rankings.
2026-09-18 ~ 2026-09-19 · 2 episodes · 20 posts
Episode 1 · Epoch AI's Benchmark Reviews: Only 4 of 15 AI Benchmarks Pass Audit, 9 Found Flawed (2026-09-18, 17 posts)
Epoch AI announced on September 18 the launch of its Benchmark Reviews program, kicking off a systematic audit of AI benchmarks. The first batch covered 15 popular benchmarks, with the conclusion that only 4 earned "Verified" status, 9 were judged flawed, and 2 could not be fully assessed due to insufficient information. The results show that most of the mainstream benchmarks widely cited across the industry failed the audit, posing a direct challenge to the credibility of model capability evaluations.
Confirmed
- Epoch AI launched the Benchmark Reviews program to systematically audit AI benchmarks
- The first batch covered 15 benchmarks: 4 Verified, 9 Flawed, and 2 unassessable due to insufficient information
- The audit found specific issues in some benchmarks — for example, 46% of sampled questions in Humanity's Last Exam (HLE) were found to have problems
- Terminal Bench 4.0 is also among the benchmarks flagged with issues
Why it matters
- Benchmark scores underpin model capability evaluations and leaderboard rankings, so most mainstream benchmarks failing the audit casts doubt on the credibility of many current model comparisons
- The program gives the community a public reference for distinguishing trustworthy from untrustworthy benchmarks, potentially driving improvements in benchmark design
- The disclosure that widely used benchmarks like HLE contain a high proportion of problematic questions is a reminder to pay attention to data quality when citing scores
- Epoch AI launches Benchmark Reviews: only 4 of first 15 benchmarks earn Verified status — xeophon · 2026-09-18
- Epoch AI launches Benchmark Reviews, audits 15 benchmarks: 4 verified, 9 flawed — Jsevillamol · 2026-09-18
- Epoch AI launches Benchmark Reviews, auditing 15 benchmarks — 9 found flawed — Jsevillamol · 2026-09-18
- Epoch AI audits 15 AI benchmarks, finds 9 flawed — 46% of sampled HLE questions broken — morqon · 2026-09-18
- Epoch launches Benchmark Reviews, auditing 15 AI benchmarks: only 4 verified, 9 flawed — stochasticchasm · 2026-09-18
- Epoch AI audits 15 AI benchmarks: only 4 pass, 9 found flawed — iamrobotbear · 2026-09-18
- Audit finds 9 of 15 AI benchmarks flawed: 45.5% of Terminal Bench 4.0 tasks broken — ajratner · 2026-09-18
- Epoch AI audits 15 AI benchmarks: 4 Verified, 9 Flawed in new initiative — pvncher · 2026-09-18
- Epoch AI audits 15 AI benchmarks: only 4 safe to trust at face value — Jsevillamol · 2026-09-18
- Epoch AI audits AI benchmarks: 9 of 15 reviewed benchmarks found flawed — joecole · 2026-09-18
- Epoch AI audits 15 AI benchmarks: 4 Verified, 9 Flawed; PostTrainBench passes and preps v1.2 — maksym_andr · 2026-09-18
- Epoch AI audits 15 AI benchmarks: 9 flawed, only 4 verified — Jsevillamol · 2026-09-19
- Epoch AI Launches Benchmark Reviews, Finds 9 of 15 Agent Benchmarks Flawed — burny_tech · 2026-09-19
- Epoch AI's benchmark audit calls 9 of 15 flawed; TB4 authors push back — DimitrisPapail · 2026-09-19
- Epoch AI audits 15 benchmarks; Anthropic says 40% of Critical Point physics questions are broken — Jsevillamol · 2026-09-19
- Epoch AI finds 46% of audited Humanity's Last Exam questions have accuracy-altering errors — charles_irl · 2026-09-19
- Epoch AI audits 15 AI benchmarks — 4 Verified, 9 Flawed — but the approach draws fire — simonguozirui · 2026-09-19
Episode 2 · Audit Sparks Debate Over Terminal Bench 4.0 Flaws (2026-09-19, 3 posts)
An audit flagged 45.5% of Terminal Bench 4.0 tasks as broken, and Epoch found scoring flaws in 30 of 66 tasks, but maintainer AlexGDimakis countered that known issues affect less than 3% of the leaderboard, fueling debate over the benchmark's integrity.
- Audit flags 45.5% of Terminal Bench 4.0 tasks as broken; maintainers push back on rigor — DimitrisPapail · 2026-09-19
- Terminal-Bench 4.0 scoring flaws hit <3% of rollouts, despite Epoch finding 30/66 defective tasks — DimitrisPapail · 2026-09-19
- Terminal Bench audit backlash: known issues affect under 3% of leaderboard, researcher says — AlexGDimakis · 2026-09-19