Audit finds up to 12% of GPQA, MMLU-Pro, and MMMU-Pro questions were broken

pawofdoom · reddit · 2026-07-29

Audit finds up to 12% of GPQA, MMLU-Pro, and MMMU-Pro questions were broken

A Reddit post points to a paper and repository auditing several popular benchmarks for malformed questions, wrong answer keys, and multiple plausible answers. The audit found that about 12% of questions in GPQA-Extended, MMLU-Pro, and MMMU-Pro were verifiably broken.

After removing the bad items, top models reportedly reached around 98% on the cleaned versions, up from the low 90s on the originals. The author also released -Clean benchmark versions, a full ledger of flagged items, dual original-vs-clean scoring, lm-eval-harness tasks, and Hugging Face datasets.

Related event: Study Reveals Heavy Contamination in Major LLM Benchmarks(2 posts)→

Original post →

More from Research

Research channel →