Study Reveals 12% of Questions in Major LLM Benchmarks Are Broken, Skewing Scores
pawofdoom · reddit · 2026-07-29
A researcher auditing mainstream LLM benchmarks discovered significant data pollution in evaluations like GPQA, MMLU-Pro, and MMMU-Pro.
- Alarming Error Rate: Approximately 12% of the questions across these benchmarks were verifiably broken, featuring malformed questions, incorrect answer keys, or multiple realistic answers.
- True Model Performance: Top models were previously topping out at 92-93% on GPQA-Diamond, which the author suspected was due to these flawed questions. When evaluated on the fixed -Clean versions of the benchmarks, top models scored around 98%.
- Open Resources: The author has released a full paper alongside -Clean versions of the benchmarks, complete with a flagged-candidate ledger, lm-eval-harness tasks, and Hugging Face datasets.
Related event: Study Reveals Heavy Contamination in Major LLM Benchmarks(2 posts)→
More from Models
- A post claims GPT-5 is now worse than a laptop-runnable Qwen3.6-27B — iamaliveix · 2026-07-29
- Users discuss what they actually use Opus 5 for beyond coding — remilouf · 2026-07-29
- Anthropic Hints at Achieving Recursive Self-Improvement, Calls for Pacing Frontier — daniel_mac8 · 2026-07-29
- Claude Opus 5 tops DeepSWE with a 74% score and a claimed 28% cost edge — daniel_mac8 · 2026-07-29
- LiquidAI’s 230M LFM2.5 encoder trends on Hugging Face — LiquidAI · 2026-07-29
- Anthropic may be 1.5 generations ahead internally, with Fable 5.1 weeks away — haider1 · 2026-07-29