MMLU-Redux paper finds ~6.5% of MMLU questions contain labeling errors

PMinervini · x · 2026-09-19

PMinervini points Epoch AI to the MMLU-Redux paper (arXiv:2406.04127), where the team manually re-annotated MMLU questions and found widespread ground-truth errors:

The authors call for revising MMLU's error-ridden questions to restore benchmark reliability.

Related event: Study Finds ~6.5% of MMLU Benchmark Questions Mislabelled(2 posts)→

Original post →

More from Models

Models channel →