Epoch AI audit finds 46% of sampled Humanity's Last Exam questions are defective
geoffwolfe · x · 2026-09-24
Epoch AI audited a random sample of 48 Humanity's Last Exam questions (6 from each of 8 categories) and designated the benchmark "Flawed": 22 (46%) had substantial accuracy-altering errors, including 12 impossible to answer correctly as written (missing assumptions, contradictory premises, subjective questions demanding objective answers). Errors were surfaced with Fable 5 and manually verified. Example: a CS question asking for a 4-point DFT over an 8-entry sequence. The upshot: part of the remaining model-vs-benchmark gap reflects the test itself, not the models.
More from Research
- Photonic matrix core on thin-film lithium niobate runs in-situ backpropagation at 8-bit precision — jwt0625 · 2026-09-25
- Agentic robotics keeps bolting GPT-6 onto harnesses — but where are systems that learn from physical failure? — DJiafei · 2026-09-25
- Dev Proposes "Evidentiality Framework": Provenance Markers for Every LLM Claim — jzesbaugh · 2026-09-25
- New Paper Quantifies When Rubric-Based RL Genuinely Improves Models vs. Hacks the Verifier — heghbalz · 2026-09-25
- DeepSeek's DSec sandbox platform serves 3M sandboxes daily for agent RL training; Liang Wenfeng co-authors — teortaxesTex · 2026-09-25
- Protein design team secures tons of GPUs and lab equipment in single-day buildout — nlarusstone · 2026-09-25