Data Duplication Wastes 33% Compute: ICML Paper Reveals Scaling Pitfalls
RylanSchaeffer · x · 2026-08-12
In their paper Scale Dependent Data Duplication, Rylan Schaeffer and colleagues explore how data repetition during pretraining affects model scaling laws, attempting to predict macroscopic data effectiveness from microscopic mechanisms.
Key insights and findings:
- Repetition is scale-dependent: As model parameters and training tokens increase, what qualifies as a "duplicate" changes. Larger models can "recognize" semantically similar but superficially different documents (e.g., translations) as repeats.
- Semantic collisions: When a corpus scales to hundreds of billions of tokens, nearest-neighbor similarities deviate sharply from an isotropic power law baseline, indicating accelerated semantic collisions.
- Larger models suffer more: When training on data sampled with replacement from a finite pool of unique documents, small models experience mild degradation, while larger models face rapidly increasing loss penalties, breaking naive scaling extrapolations. Previous research suggested that data duplication wastes about 33% of compute.
Related event: ICML Paper Reveals Data Duplication Wastes 33% of Compute(2 posts)→
More from Research
- Recent Humanoid Robotics Papers: Breakthroughs in Parkour and Single-Leg Balance — carlosdponx · 2026-08-12
- AlphaFold Has Produced Zero Approved Drugs So Far, Investor Notes — JosephJacks_ · 2026-08-12
- DMSampler Accelerates Diffusion RL Training, Cutting GPU Hours by 10x — jiqizhixin · 2026-08-12
- RLHF Book officially published; Nathan Lambert moves to independent research — Stefania_druga · 2026-08-12
- Co-Arena: live arena for computer-use agents hits 55K steps in 6 days — Scobleizer · 2026-08-12
- The Math Proves It: Why AI Agents Are Not 'Digital Humans' — Independent-Key-1621 · 2026-08-12