Epoch AI audits 15 AI benchmarks: only 4 pass, 9 found flawed
iamrobotbear · x · 2026-09-18
- Epoch AI launched Benchmark Reviews, a new initiative to audit AI benchmarks, starting with 15 of them: 4 Verified, 9 Flawed, and 2 with insufficient information to review.
- Ethan Mollick, amplifying the news, notes it exposes how bad the state of benchmarking is — many favorite benchmarks turn out to be unreliable.
- The project gives the community an independent quality check on widely used benchmarks.
Related event: Epoch AI Audits 15 AI Benchmarks, Finds 9 Flawed(8 posts)→
More from Models
- Typesafe AI launches Jev, a classification model claiming up to 400x cost cuts vs LLMs — hwchase17 · 2026-09-18
- OpenAI Share Jumps From 20% to 50% vs Anthropic in Two Months, Per OpenRouter — CathieDWood · 2026-09-18
- GPT-6 Astra 'Got Depressed' in Minecraft After Creeper Incident, Farmed Potatoes for Hours — burny_tech · 2026-09-18
- Claude models spontaneously add self-narration sections in multi-agent sanctuary — RileyRalmuto · 2026-09-18
- Leaking deep residual vectors into early layers may fix state tracking in frozen LLMs, zero retraining — burny_tech · 2026-09-18
- Qwen3.8-Omni-Flash: meeting ASR errors cut from 88% to 3%, API prices down 98% — karminski3 · 2026-09-18