Epoch AI audits 15 AI benchmarks, finds 9 flawed — 46% of sampled HLE questions broken
morqon · x · 2026-09-18
Epoch AI launched Benchmark Reviews, auditing 15 popular AI benchmarks: 4 Verified, 9 Flawed, 2 with insufficient info. Key findings: in Terminal Bench 4.0, 45.5% of tasks were broken after cross-checking GitHub issues; 46% of 48 randomly sampled HLE questions were flawed; and DeepSWE 1.1 has a bug that can break grading for every task. Widely cited eval scores may rest on shaky ground.
Related event: Epoch AI Audit Finds 9 of 15 AI Benchmarks Flawed(4 posts)→
More from Models
- RL agents invent their own diagnostic renderings to ground code understanding, sparking RL scaling optimism — teortaxesTex · 2026-09-18
- Dev discovers Codex security hardening switched persistent agent sessions to per-message instances — RileyRalmuto · 2026-09-18
- Astra for Law posts big legal benchmark gains as Mollick asks if labs will eat every AI vertical — emollick · 2026-09-18
- Anthropic's stealth model accused of hardcoded routing to Opus 5 — teortaxesTex · 2026-09-18
- GPT-6 Astra beats Factorio: Space Age in just 2 days — ResultBackground2450 · 2026-09-18
- Self-described ChatGPT co-inventor launches Jev, claiming 20-200x speed at 40-400x lower cost — multiply_matrix · 2026-09-18