Epoch AI audits 15 AI benchmarks: 9 flawed, only 4 verified
Jsevillamol · x · 2026-09-19
Epoch AI has launched Benchmark Reviews, a new initiative to audit AI evaluation benchmarks, starting with 15 benchmarks: 4 Verified, 9 Flawed, and 2 lacking sufficient information for review. The project addresses growing concerns about benchmark quality and the reliability of model evaluations.
Related event: Epoch AI's Benchmark Reviews: Only 4 of 15 AI Benchmarks Pass Audit(12 posts)→
More from Research
- AM-Bench: A Simulation Suite and Benchmark for Aerial Manipulation — GuanyaShi · 2026-09-19
- First benchmark to compare official Jev against open-source lookalikes in the works — airesearch12 · 2026-09-19
- Study finds scientific beauty and impact are surprisingly correlated — jacobkimmel · 2026-09-19
- Trolley problem: Jev-style API turns jina-reranker-v3.5 into a ruthless decision engine — gaganghotra_ · 2026-09-19
- Probe guidance steers continuous diffusion LMs with just 1-3% extra inference compute — itsbautistam · 2026-09-19
- LLM Analysis of Wikipedia Battle Pages Crowns Napoleon the GOAT With 16.7 Career WAR — ctjlewis · 2026-09-19