Synthetic Pre-pretraining Survives Scaling to 7B, but Gains Come From Long-range Retrieval, Not Grammar

nikaletras · x · 2026-10-01

The largest-ever study of pre-pretraining (PPT) — briefly training on synthetic non-natural-language data before pretraining — is now on arXiv, spanning 5 PPT tasks, 4 pretraining data mixtures, model scales from 500M to 7B, and budgets up to 100B tokens.

Key findings:

The authors frame PPT as a low-cost addition to pretraining.

Related event: Synthetic pre-pretraining scales to 7B models, saving 21B tokens(2 posts)→

Original post →

More from Research

Research channel →