Data Duplication Can Waste Up to 33% of LLM Pretraining Compute

RylanSchaeffer · x · 2026-07-05

It is widely recognized that data duplication harms LLM pretraining. A paper by Chudnovsky further demonstrates that this damage depends on a predictable scaling interaction among model parameter count, the number of duplicated documents, and duplication frequency. A poor combination can severely waste computing power, with losses reaching up to approximately 33%.

Related event: Stanford Study: Data Duplication Can Waste 33% of LLM Pretraining Compute(2 posts)→

Original post →

More from Research

Research channel →