Stanford Study: Pre-training Data Repetition Can Waste 33% of Compute
sanmikoyejo · x · 2026-07-05
A Stanford team (jchudnov; co-first authors Joshua K. and Noam Levi; along with Rylan Schaeffer, Yegor D., Bo He, Sanmi Koyejo, etc.; affiliated with Stanford HAI / AI Lab / NLP) released new research on pre-training data repetition.
As pre-training becomes increasingly constrained by data volume, data repetition is inevitable. However, the cost of repetition cannot be measured solely by the "proportion of repeated data"—datasets with the same proportion can have vastly different risks. The study found that the harm of repetition peaks at "moderate repetition counts": small data pools repeated numerous times are memorized quickly, while large data pools repeated a few times are memorized more slowly but are equally harmful. Furthermore, the peak position shows a clear trend with model scale: larger models peak later, a pattern that is largely independent of architecture.
In the worst case, residual repetition can waste about 33% of compute power, which is hard to detect just by looking at loss curves. Although deduplication is standard practice, it is not perfect. The team used a new method to quantify the actual harm of residual repetition.
Related event: Stanford Study: Data Duplication Can Waste 33% of LLM Pretraining Compute(2 posts)→
More from Research
- A quick QAT turned into a much longer grind than expected — cephaloform · 2026-07-27
- Microsoft’s ReOPD reuses teacher prefixes to make agent distillation 4× faster — dair_ai · 2026-07-27
- A study of 12,750 arXiv papers finds AI-like writing flagged in 65% of CS papers — 新智元 · 2026-07-27
- Yaqi Xie joins UIUC as assistant professor and starts recruiting for AI agents and robots — dhruv2038 · 2026-07-27
- Kimi K3 may be strong on cyber, but token efficiency keeps it off UK AISIS — teortaxesTex · 2026-07-27
- ARC AGI 3 should have stayed private, with no examples or public dataset — flowersslop · 2026-07-27