Epoch AI Launches Benchmark Reviews, Finds 9 of 15 Agent Benchmarks Flawed
burny_tech · x · 2026-09-19
- Epoch AI launched Benchmark Reviews, a new initiative to audit AI benchmarks, starting with 15 AI agent benchmarks.
- Initial results: only 4 Verified, 9 Flawed, and 2 with insufficient information to review.
- ddkang echoed the news, pointing to his team's July 2025 work establishing best practices for rigorous agent benchmarks — long-standing concerns now validated by the audit.
More from Models
- Sentdex tries DeepSeek v4.1f, says he feels like he's 'cheating on GLM' — Sentdex · 2026-09-19
- Ternary Bonsai 2 27B at 1.75bpw fits an 8GB GPU, hits 93.3% accuracy in audiobook speaker-attribution test — autonoma_2042 · 2026-09-19
- Google finally has a SOTA rogue model, and the AI crowd is joking about relief — rao2z · 2026-09-19
- ChatGPT keeps answering how/why questions with a leading Yes, Reddit user finds — Starrryyyyy · 2026-09-19
- One Model, Three Skills: Programmatic Use, Chat, and Test-Taking Diverge — lateinteraction · 2026-09-19
- Want US frontier lab secrets? Just look at Chinese SOTA models, quips AI insider — gowthami_s · 2026-09-19