Audit finds up to 12% of GPQA, MMLU-Pro, and MMMU-Pro questions were broken
pawofdoom · reddit · 2026-07-29
Audit finds up to 12% of GPQA, MMLU-Pro, and MMMU-Pro questions were broken
A Reddit post points to a paper and repository auditing several popular benchmarks for malformed questions, wrong answer keys, and multiple plausible answers. The audit found that about 12% of questions in GPQA-Extended, MMLU-Pro, and MMMU-Pro were verifiably broken.
After removing the bad items, top models reportedly reached around 98% on the cleaned versions, up from the low 90s on the originals. The author also released -Clean benchmark versions, a full ledger of flagged items, dual original-vs-clean scoring, lm-eval-harness tasks, and Hugging Face datasets.
Related event: Study Reveals Heavy Contamination in Major LLM Benchmarks(2 posts)→
More from Research
- Discussion: Why hasn't anyone built a neural network to detect AI text? Image detection has research papers — emeka_boris · 2026-07-30
- Agents still struggle with mathematical work: Codex spirals into 'proof certificates' and inventories — doodlestein · 2026-07-30
- TorchSpec Enables Disaggregated Speculative Decoding Training at Scale — zhyncs42 · 2026-07-30
- Compute Surge: 10 Major Scientific Breakthroughs AI Could Unlock by 2028 — Annual_Judge_7272 · 2026-07-30
- Inside SOTA Deep Research: Native Model Training and 150 Sub-Agents — SimonShaoleiDu · 2026-07-30
- Top AI Startups Are Barely Publishing Their Research Anymore — YeGoblynQueenne · 2026-07-30