GPQA Benchmark Suspected of Data Contamination

teortaxesTex · x · 2026-07-16

This post highlights a frustrating research phenomenon: someone took the GPQA Diamond evaluation data, made "very minor rewrites," and trained on it for 10 epochs.

The main takeaway is that even light rewrites can severely contaminate benchmark results if the training and evaluation sets overlap. This inflates the model's perceived performance, making the scores unreliable.

Related event: SOOFI benchmark claims challenged over leakage and baseline reporting(13 posts)→

Original post →

More from Research

Research channel →