Developer Pushes Back: Humanity's Last Exam Dataset Has Very Low Error Rate
Bayesian0_0 · x · 2026-09-02
In response to claims that the Humanity's Last Exam dataset contains many errors and suffers from poor methodology, developer Sauers pushed back, saying he personally checked many of the biology questions and found possibly only one that was debatable. He argued that the accuracy statistics are informative and that the related analysis article should not simply be deleted.
More from Research
- Boaz Barak: abandoning chain-of-thought before validated alternatives is irresponsible — inductionheads · 2026-09-03
- Developer once tried building AI benchmark from Puzzlescript, similar to ARC-AGI-3 — Darpinian · 2026-09-03
- He quarantined pre-1996 sources to build a 'clone' of Prof. Milhaupt as a sounding board — KarlMuth · 2026-09-03
- Do induction heads already explain LLMs' 'unprecedented' abilities? Researchers debate — aryaman2020 · 2026-09-03
- Do induction heads and attention sinks count? Debate over interpretability's missed milestone — aryaman2020 · 2026-09-03
- Counterfactual debugging scales sim2real failure diagnosis to 1M steps in world models — sarahcat21 · 2026-09-03