Study Reveals 12% of Questions in Major LLM Benchmarks Are Broken, Skewing Scores

pawofdoom · reddit · 2026-07-29

A researcher auditing mainstream LLM benchmarks discovered significant data pollution in evaluations like GPQA, MMLU-Pro, and MMMU-Pro.

Related event: Study Reveals Heavy Contamination in Major LLM Benchmarks(2 posts)→

Original post →

More from Models

Models channel →