Epoch AI audits 15 AI benchmarks — 4 Verified, 9 Flawed — but the approach draws fire

simonguozirui · x · 2026-09-19

Epoch AI launched Benchmark Reviews, an initiative to audit AI benchmarks, starting with 15: 4 Verified, 9 Flawed, and 2 lacking enough information. Reposting it, alexgshaw argues the binary "flawed vs verified" labeling is flawed: benchmarks are software and all software has bugs. He advocates continuous benchmarks — giving anyone tools to create, improve, and maintain benchmarks while reconciling results to the latest version — and invites Epoch to encode its verification practices into such tools.

Related event: Epoch AI's Benchmark Reviews: Only 4 of 15 AI Benchmarks Pass Audit, 9 Found Flawed(17 posts)→

Original post →

More from Models

Models channel →