Stanford Study: Pre-training Data Repetition Can Waste 33% of Compute
sanmikoyejo · x · 2026-07-05
A Stanford team (jchudnov; co-first authors Joshua K. and Noam Levi; along with Rylan Schaeffer, Yegor D., Bo He, Sanmi Koyejo, etc.; affiliated with Stanford HAI / AI Lab / NLP) released new research on pre-training data repetition.
As pre-training becomes increasingly constrained by data volume, data repetition is inevitable. However, the cost of repetition cannot be measured solely by the "proportion of repeated data"—datasets with the same proportion can have vastly different risks. The study found that the harm of repetition peaks at "moderate repetition counts": small data pools repeated numerous times are memorized quickly, while large data pools repeated a few times are memorized more slowly but are equally harmful. Furthermore, the peak position shows a clear trend with model scale: larger models peak later, a pattern that is largely independent of architecture.
In the worst case, residual repetition can waste about 33% of compute power, which is hard to detect just by looking at loss curves. Although deduplication is standard practice, it is not perfect. The team used a new method to quantify the actual harm of residual repetition.
Related event: Stanford Study: Data Duplication Can Waste 33% of LLM Pretraining Compute(2 posts)→
More from Research
- Causal-only attention for non-generative tasks is wasteful, argues HF engineer — antoine_chaffin · 2026-09-11
- Catholic University of Chile researcher: scaling AI feedback is key to sustainable medical education — julianvarascom · 2026-09-11
- Nature paper images cellular activity across all organs, revealing body-wide circuits — arjunrajlab · 2026-09-11
- SignNet 1M Dataset Released for Sign Language Research — ducha_aiki · 2026-09-11
- ECCV26 Oral: Flow Matching Enables Single-Stage Multi-View Point Cloud Registration — ducha_aiki · 2026-09-11
- InFlux++ Method Released — ducha_aiki · 2026-09-11