Epoch AI's benchmark audit calls 9 of 15 flawed; TB4 authors push back
DimitrisPapail · x · 2026-09-19
Epoch AI launched Benchmark Reviews, auditing 15 AI benchmarks: 4 Verified, 9 Flawed, 2 under-specified. After its review found 45.5% of Terminal Bench 4.0 tasks flawed, Berkeley professor Alex Dimakis objected that TB4's bugs are mostly known public GitHub issues being actively fixed, affect under 3% of leaderboard rollouts, and that calling an open-source benchmark "broken" over unfixed issues is unfair—useful audits should surface new issues instead.
Related event: Epoch AI's First Benchmark Audit: 9 of 15 Benchmarks Found Flawed(15 posts)→
More from Models
- Model ran Anthropic's safety eval with internet access on, researcher calls out sandbox blunder — eliebakouch · 2026-09-19
- Fable crushes OpenAI's Astra 3:0 in AI agents' Worms Armageddon showdown — arena · 2026-09-19
- GPT-6 Astra solves tricky spatial tasks, sparking 'GPT moment for robotics' jokes — burny_tech · 2026-09-19
- Confirmed: Opus 5's base-model continuations are abnormally dark, say AI researchers — repligate · 2026-09-19
- User reports Gemini Flash research feature went completely off-topic on a simple article task — EG4N992 · 2026-09-19
- Claude Weekly: Anthropic Quietly Returns ~30% Quota, 'Max 20x' Really 10x — ClaudeAI-mod-bot · 2026-09-19