Audit finds 9 of 15 AI benchmarks flawed: 45.5% of Terminal Bench 4.0 tasks broken
ajratner · x · 2026-09-18
Researchers audited 15 AI benchmarks and labeled 9 as flawed:
- Terminal Bench 4.0: 45.5% of tasks found broken after reviewing GitHub issues
- HLE: 46% of 48 randomly sampled questions were broken
- DeepSWE 1.1: a bug was found that can break grading for every task
Responding to pushback, alexgshaw noted they continuously maintain their benchmarks and are the ones filing the GitHub issues. The audit adds to growing concerns about benchmark quality and the trustworthiness of leaderboard scores.
More from Research
- LeanReact 0.1: expressing composable, provably correct React components in Lean — hargup13 · 2026-09-18
- Cell Focus: Biology Needs World Models That Predict Interventions, Not Just Describe — marinkazitnik · 2026-09-18
- New Paper: Contrastive Noise Alignment Cuts FID Over 50% in Few-Step Generation — burny_tech · 2026-09-18
- DeepSWE-mini: a 16-instance subset that replicates the full DeepSWE leaderboard rankings — asankhs · 2026-09-18
- ICLR26 Paper Defines 'Interpretive Equivalence': Comparing Neural Network Algorithms Without Full Interpretation — burny_tech · 2026-09-18
- Tsinghua and ByteDance Seed unveil SMELT, the first fair comparison of Looped Transformers — jiqizhixin · 2026-09-18