Epoch AI audits 15 AI benchmarks, finds 9 flawed — 46% of sampled HLE questions broken

morqon · x · 2026-09-18

Epoch AI launched Benchmark Reviews, auditing 15 popular AI benchmarks: 4 Verified, 9 Flawed, 2 with insufficient info. Key findings: in Terminal Bench 4.0, 45.5% of tasks were broken after cross-checking GitHub issues; 46% of 48 randomly sampled HLE questions were flawed; and DeepSWE 1.1 has a bug that can break grading for every task. Widely cited eval scores may rest on shaky ground.

Related event: Epoch AI Audit Finds 9 of 15 AI Benchmarks Flawed(4 posts)→

Original post →

More from Models

Models channel →