Audit finds 9 of 15 AI benchmarks flawed: 45.5% of Terminal Bench 4.0 tasks broken

ajratner · x · 2026-09-18

Researchers audited 15 AI benchmarks and labeled 9 as flawed:

Responding to pushback, alexgshaw noted they continuously maintain their benchmarks and are the ones filing the GitHub issues. The audit adds to growing concerns about benchmark quality and the trustworthiness of leaderboard scores.

Related event: Epoch AI Launches Benchmark Reviews: Only 4 of 15 Popular AI Benchmarks Verified(8 posts)→

Original post →

More from Research

Research channel →