MMLU-Redux paper finds ~6.5% of MMLU questions contain labeling errors
PMinervini · x · 2026-09-19
PMinervini points Epoch AI to the MMLU-Redux paper (arXiv:2406.04127), where the team manually re-annotated MMLU questions and found widespread ground-truth errors:
- An estimated 6.49% of MMLU questions contain errors; the Virology subset shows a 57% error rate
- They propose a systematic error annotation protocol and release MMLU-Redux, 5,700 manually re-annotated questions across all 57 subjects
- Re-evaluating on MMLU-Redux reveals significant discrepancies with originally reported model scores
The authors call for revising MMLU's error-ridden questions to restore benchmark reliability.
Related event: Study Finds ~6.5% of MMLU Benchmark Questions Mislabelled(2 posts)→
More from Models
- Kimi's cryptic post decodes to pi, hinting at imminent Kimi K3.1 release (unconfirmed) — kimmonismus · 2026-09-19
- Naming matters: Jev's traction and BERT's success owe much to their names — IgorCarron · 2026-09-19
- Cursor Ultra perks via SuperGrok Heavy quietly cut 75%: $400 to $100 quota — mazzaTalk · 2026-09-19
- 4x RTX 3090 advice: keep Qwen 27B Q8 or switch to Qwen Next Flash — Zyj · 2026-09-19
- Real enterprise build shows popular LLM benchmarks are largely irrelevant — DivideHorror3217 · 2026-09-19
- Nearly 20 openjev models cataloged as hobbyist preps first community leaderboard — airesearch12 · 2026-09-19