Epoch AI launches Benchmark Reviews, auditing 15 benchmarks — 9 found flawed
Jsevillamol · x · 2026-09-18
Epoch AI launched Benchmark Reviews, a new initiative to systematically audit AI benchmarks. The first batch covers 15 benchmarks: 4 Verified, 9 Flawed, and 2 with not enough information for a review.
Lead Alex Barry says people deep in AI benchmarking develop a sense of which benchmarks to trust, and the project aims to make that knowledge accessible while raising the overall quality of benchmarks over time.
Related event: Epoch AI Audit Finds 9 of 15 AI Benchmarks Flawed(4 posts)→
More from Models
- RL agents invent their own diagnostic renderings to ground code understanding, sparking RL scaling optimism — teortaxesTex · 2026-09-18
- Dev discovers Codex security hardening switched persistent agent sessions to per-message instances — RileyRalmuto · 2026-09-18
- Astra for Law posts big legal benchmark gains as Mollick asks if labs will eat every AI vertical — emollick · 2026-09-18
- Anthropic's stealth model accused of hardcoded routing to Opus 5 — teortaxesTex · 2026-09-18
- GPT-6 Astra beats Factorio: Space Age in just 2 days — ResultBackground2450 · 2026-09-18
- Self-described ChatGPT co-inventor launches Jev, claiming 20-200x speed at 40-400x lower cost — multiply_matrix · 2026-09-18