Study finds semantic dupes for 78% of CodeForces problems in training data

gleech · x · 2026-09-04

In a thread with Yoav Goldberg on how to measure LLM progress, Gavin Leech cites an arXiv paper on soft contamination: semantic duplicates in training data that n-gram filters miss. Embedding the Olmo3 corpus, authors found duplicates for 78% of CodeForces and exact dupes for 50% of ZebraLogic problems; adding benchmark dupes to training improves scores, and finetuning on dupes also lifts truly held-out scores from the same benchmark. Recent benchmark gains are thus confounded. Proposed fixes: preregistered forecasting-style task streams and a benchmark treadmill with quarterly refreshes.

Related event: Yoav Goldberg: Any Scorable Benchmark Will Eventually Be Brute-Forced(3 posts)→

Original post →

More from Research

Research channel →