Epoch AI audit finds 46% of sampled Humanity's Last Exam questions are defective

geoffwolfe · x · 2026-09-24

Epoch AI audited a random sample of 48 Humanity's Last Exam questions (6 from each of 8 categories) and designated the benchmark "Flawed": 22 (46%) had substantial accuracy-altering errors, including 12 impossible to answer correctly as written (missing assumptions, contradictory premises, subjective questions demanding objective answers). Errors were surfaced with Fable 5 and manually verified. Example: a CS question asking for a 4-point DFT over an 8-entry sequence. The upshot: part of the remaining model-vs-benchmark gap reflects the test itself, not the models.

Original post →

More from Research

Research channel →