Epoch AI audits 15 AI benchmarks: 4 Verified, 9 Flawed in new initiative

pvncher · x · 2026-09-18

Epoch AI launched Benchmark Reviews, a new initiative to audit AI benchmarks, starting with 15 of them: 4 Verified, 9 Flawed, and 2 with insufficient information for review. Quoting @charliermarsh: in DeepSWE v1.1, the verifier discards the agent's changes to some test files without the model knowing, causing various 'irrelevant' failures.

Related event: Epoch AI Launches Benchmark Reviews: Only 4 of 15 Popular AI Benchmarks Verified(8 posts)→

Original post →

More from Models

Models channel →