Stanford Study: Pre-training Data Repetition Can Waste 33% of Compute

sanmikoyejo · x · 2026-07-05

A Stanford team (jchudnov; co-first authors Joshua K. and Noam Levi; along with Rylan Schaeffer, Yegor D., Bo He, Sanmi Koyejo, etc.; affiliated with Stanford HAI / AI Lab / NLP) released new research on pre-training data repetition.

As pre-training becomes increasingly constrained by data volume, data repetition is inevitable. However, the cost of repetition cannot be measured solely by the "proportion of repeated data"—datasets with the same proportion can have vastly different risks. The study found that the harm of repetition peaks at "moderate repetition counts": small data pools repeated numerous times are memorized quickly, while large data pools repeated a few times are memorized more slowly but are equally harmful. Furthermore, the peak position shows a clear trend with model scale: larger models peak later, a pattern that is largely independent of architecture.

In the worst case, residual repetition can waste about 33% of compute power, which is hard to detect just by looking at loss curves. Although deduplication is standard practice, it is not perfect. The team used a new method to quantify the actual harm of residual repetition.

Related event: Stanford Study: Data Duplication Can Waste 33% of LLM Pretraining Compute(2 posts)→

Original post →

More from Research

Research channel →