Study finds semantic dupes for 78% of CodeForces problems in training data
gleech · x · 2026-09-04
In a thread with Yoav Goldberg on how to measure LLM progress, Gavin Leech cites an arXiv paper on soft contamination: semantic duplicates in training data that n-gram filters miss. Embedding the Olmo3 corpus, authors found duplicates for 78% of CodeForces and exact dupes for 50% of ZebraLogic problems; adding benchmark dupes to training improves scores, and finetuning on dupes also lifts truly held-out scores from the same benchmark. Recent benchmark gains are thus confounded. Proposed fixes: preregistered forecasting-style task streams and a benchmark treadmill with quarterly refreshes.
Related event: Yoav Goldberg: Any Scorable Benchmark Will Eventually Be Brute-Forced(3 posts)→
More from Research
- Researcher teases dynamic composite eval index as "evals run on Twitter vibes" — evijit · 2026-09-04
- Turning agent traces into training data: capture, sampling, and labelling, worked through — spilldahill · 2026-09-04
- Researcher Visualizes Qwen 2.5 Embedding Space, Says Mainstream Models' Conversational Ability Is Flattening — arianaram · 2026-09-04
- DeepMind's Prateek Jain on MatFormer: agents that dial compute up or down by task difficulty — jainprateek_ · 2026-09-04
- Fei-Fei Li on World Labs' Atlas: new view prediction as a world model primitive — a16z Podcast · 2026-09-04
- Embodied AI dataset ACE-Data-0 hits HF trending with ~30K downloads one week after release — liuziwei7 · 2026-09-04