Paper Reveals Larger LLMs Tolerate More Data Repetition During Pretraining

heghbalz · x · 2026-08-22

This paper investigates the strategy of repeating high-quality domain data to maintain a fixed tokens-per-parameter (TPP) ratio during LLM pretraining when such data is scarce. Key findings include:

This offers a new engineering perspective on mitigating high-quality data dilution.

Related event: ICML Paper: Larger LLMs Tolerate More Data Repetition in Pretraining(2 posts)→

Original post →

More from Infra

Infra channel →