Epoch AI finds 46% of audited Humanity's Last Exam questions have accuracy-altering errors
charles_irl · x · 2026-09-19
Epoch AI's benchmark review delivers a "Flawed" verdict for Humanity's Last Exam, the knowledge benchmark created by the Center for AI Safety, Scale AI, and independent contributors.
- Method: sampled 48 questions (6 each from 8 categories), used Fable 5 to surface errors, then manually audited each against an evidence bar.
- Result: 22 of 48 questions (46%) had substantial accuracy-altering errors.
- Breakdown: 12 questions impossible to answer as written (missing assumptions, contradictory premises, multiple valid interpretations, objective answers demanded for subjective questions); 10 well-posed questions could produce false negatives, 5 of which could also produce false positives.
- Example: a CS question asks for a 4-point DFT over a sequence with 8 samples.
- The review covers the original benchmark, not HLE-Rolling or HLE-Verified.
More from Models
- Dev runs 16-model eval: Jev and Haiku tie at 0.121/0.122 on calibration error — AlexKim · 2026-09-19
- 16-model eval finds Sonnet 5 hedges most, landing in ambiguous zone 41.3% of the time — AlexKim · 2026-09-19
- Dev tests 16 models to evaluate TypeSafe's Jev — it ranked 10th on accuracy — AlexKim · 2026-09-19
- Mystery Model Jev Launches Claiming 200x Speed and 400x Cost Cuts, Devs Impressed — multiply_matrix · 2026-09-19
- GLM 5.3 Flash leads quality, Qwen 3.8 Flash Next wins speed in open small-model comparison — HankYeomans · 2026-09-19
- Ternary Bonsai 2 27B quantized to a 7GB single file, runs on an 8GB GPU — cephaloform · 2026-09-19