Internal Data Repetition Wastes 33% of Compute, ICML Paper Reveals Pretraining Hidden Costs
RylanSchaeffer · x · 2026-08-12
A new study, Internal Data Repetition Destroys Language Models, investigates the damage caused by repeated data in LLM pre-training. The research shows that even aggressively deduplicated corpora retain repetition, leading to systematic compute waste.
- The Worst-Case Scenario: Holding compute constant, repeating a moderately sized subset a moderate number of times is more damaging than repeating a large subset a few times or a small subset many times.
- Massive Compute Loss: When repeated documents consume 10% of the FLOPs budget, the compute-equivalent loss is severe. On FineWeb-Edu-Dedup, the most damaging repeat count for a 344M-parameter model wastes an equivalent of 33% of total FLOPs.
- Scaling Law: The most damaging number of repeated data grows more quickly than compute.
Related event: ICML Paper Reveals Data Duplication Wastes 33% of Compute(2 posts)→
More from Research
- AlphaFold Has Produced Zero Approved Drugs So Far, Investor Notes — JosephJacks_ · 2026-08-12
- Weekly Humanoid Robotics Papers: Motion Tracking & Balance — carlosdponx · 2026-08-12
- DMSampler Accelerates Diffusion RL Training, Cutting GPU Hours by 10x — jiqizhixin · 2026-08-12
- RLHF Book officially published; Nathan Lambert moves to independent research — Stefania_druga · 2026-08-12
- Co-Arena: live arena for computer-use agents hits 55K steps in 6 days — Scobleizer · 2026-08-12
- The Math Proves It: Why AI Agents Are Not 'Digital Humans' — Independent-Key-1621 · 2026-08-12