MMLU Is Riddled With Errors: 6.49% of Questions Wrong, 57% in Virology
PMinervini · x · 2026-09-19
The MMLU-Redux paper (arXiv:2406.04127) by Aryo Pradipta Gema, Pasquale Minervini, et al. audits the popular MMLU benchmark and finds it riddled with ground-truth errors.
Key findings:
- An estimated 6.49% of MMLU questions contain errors; the Virology subset is worst, with 57% of analyzed questions flawed.
- The team introduces a novel error-annotation protocol and releases MMLU-Redux: 5,700 manually re-annotated questions spanning all 57 MMLU subjects.
- Re-evaluating models on MMLU-Redux reveals significant discrepancies with originally reported scores, meaning bad questions have been obscuring true LLM capabilities.
- The authors advocate revising MMLU's erroneous questions to restore benchmark reliability.
Related event: Study Finds ~6.5% of MMLU Benchmark Questions Mislabelled(2 posts)→
More from Models
- Cursor Ultra perks via SuperGrok Heavy quietly cut 75%: $400 to $100 quota — mazzaTalk · 2026-09-19
- 4x RTX 3090 advice: keep Qwen 27B Q8 or switch to Qwen Next Flash — Zyj · 2026-09-19
- Real enterprise build shows popular LLM benchmarks are largely irrelevant — DivideHorror3217 · 2026-09-19
- Nearly 20 openjev models cataloged as hobbyist preps first community leaderboard — airesearch12 · 2026-09-19
- Dev Slams New DeepSeek Model as Distilled Claude Without the Intelligence — Aryvyo · 2026-09-19
- Xiaomi's AI Persona Teased: Unified Model Xiaomi MiMo Launches Tomorrow — xiaohu · 2026-09-19