Epoch AI finds 46% of audited Humanity's Last Exam questions have accuracy-altering errors

charles_irl · x · 2026-09-19

Epoch AI's benchmark review delivers a "Flawed" verdict for Humanity's Last Exam, the knowledge benchmark created by the Center for AI Safety, Scale AI, and independent contributors.

Related event: Epoch AI's Benchmark Reviews: Only 4 of 15 AI Benchmarks Pass Audit, 9 Found Flawed(17 posts)→

Original post →

More from Models

Models channel →