Epoch AI audits 15 AI benchmarks: 4 Verified, 9 Flawed in new initiative
pvncher · x · 2026-09-18
Epoch AI launched Benchmark Reviews, a new initiative to audit AI benchmarks, starting with 15 of them: 4 Verified, 9 Flawed, and 2 with insufficient information for review. Quoting @charliermarsh: in DeepSWE v1.1, the verifier discards the agent's changes to some test files without the model knowing, causing various 'irrelevant' failures.
More from Models
- Claude Picks the Voynich Manuscript, Makes a Movie — and It Cost About $85 in Tokens — emollick · 2026-09-18
- DeepSeek appears to tighten NSFW moderation, ending its Anything-Goes era — teortaxesTex · 2026-09-18
- Rumor: xAI's Grok 4.7 release is imminent — kimmonismus · 2026-09-18
- GitLab hosts Kimi K3, GLM 5.3 and MiniMax M3, up to 4x calls per credit — RealGeneKim · 2026-09-18
- 10 hours of heavy Codex coding barely moved a 6% weekly allowance — ___Patrice___ · 2026-09-18
- ChatGPT co-inventor launches Jev: claims 20-200x faster, 40-400x cheaper than LLMs — GabGarrett · 2026-09-18