Skeptical deep dive confirms Humanity's Last Exam errors; official o3-mini grader marked right answers wrong every time
paul_cal · x · 2026-09-20
paulcal ran a skeptical deep dive into claimed issues with Humanity's Last Exam and confirmed that every question he checked was indeed wrong. In one example, he tested how often the official grader setup (an o3-mini-based HLM grader) followed the answer key — and it marked a correct answer as wrong every single time. The quoted post notes HLE appears to be 50% wrong and questions why nobody caught these errors before publication, undermining confidence in HLE-based frontier model claims.
More from Models
- XGEN debuts Generative World Simulation: JING model tops WBench Full split — hey_abusiddik · 2026-09-20
- Code-only heuristic policies can beat frontier models on Craftax, evals researcher says — JoshPurtell · 2026-09-20
- Codex usage reset now live for all, big OpenAI release teased for Tuesday — kimmonismus · 2026-09-20
- Jev beats GPT-5.6 Luna on PR review: 1.93x faster at $0.0014 per run — aniketmaurya · 2026-09-20
- Fruit fly connectome chess model beats Jev 4-1 in 10 games, with a playable demo site — maximelabonne · 2026-09-20
- Bonsai 2 27B safety guardrails reportedly cut SWE-bench and Terminal-bench scores by ~20 points — julianharris · 2026-09-20