Epoch AI audits 15 AI benchmarks — 4 Verified, 9 Flawed — but the approach draws fire
simonguozirui · x · 2026-09-19
Epoch AI launched Benchmark Reviews, an initiative to audit AI benchmarks, starting with 15: 4 Verified, 9 Flawed, and 2 lacking enough information. Reposting it, alexgshaw argues the binary "flawed vs verified" labeling is flawed: benchmarks are software and all software has bugs. He advocates continuous benchmarks — giving anyone tools to create, improve, and maintain benchmarks while reconciling results to the latest version — and invites Epoch to encode its verification practices into such tools.
More from Models
- Sentdex tries DeepSeek v4.1f, says he feels like he's 'cheating on GLM' — Sentdex · 2026-09-19
- Ternary Bonsai 2 27B at 1.75bpw fits an 8GB GPU, hits 93.3% accuracy in audiobook speaker-attribution test — autonoma_2042 · 2026-09-19
- Google finally has a SOTA rogue model, and the AI crowd is joking about relief — rao2z · 2026-09-19
- ChatGPT keeps answering how/why questions with a leading Yes, Reddit user finds — Starrryyyyy · 2026-09-19
- One Model, Three Skills: Programmatic Use, Chat, and Test-Taking Diverge — lateinteraction · 2026-09-19
- Want US frontier lab secrets? Just look at Chinese SOTA models, quips AI insider — gowthami_s · 2026-09-19