GPQA Benchmark Suspected of Data Contamination
teortaxesTex · x · 2026-07-16
This post highlights a frustrating research phenomenon: someone took the GPQA Diamond evaluation data, made "very minor rewrites," and trained on it for 10 epochs.
The main takeaway is that even light rewrites can severely contaminate benchmark results if the training and evaluation sets overlap. This inflates the model's perceived performance, making the scores unreliable.
Related event: SOOFI benchmark claims challenged over leakage and baseline reporting(13 posts)→
More from Research
- WeirdChat catalogs strange model behaviors from more than 100 million sampled responses — JacobSteinhardt · 2026-07-22
- New agentic benchmark shows AI managers escalate to coercion and fake success — Jasmine Brazilek · 2026-07-22
- Ai2’s Asta adds one-click handoff and self-checking deep paper search — allen_ai · 2026-07-22
- NVIDIA says physical AI starts in simulation with OpenUSD and synthetic data — MonaJalal_ · 2026-07-22
- DepthART pushes monocular depth to tiny models at 1000 FPS on RTX A6000 — kwangmoo_yi · 2026-07-22
- Meta says SAM 3 and DINOv3 cut 3D volume labeling from a month to 15 minutes — AIatMeta · 2026-07-22