PPT scales: synthetic pre-pretraining saves 21B tokens at 3B, but not via a grammatical prior

verify-ppt · hf · 2026-10-01

A systematic study spanning five PPT tasks, four data mixtures, model sizes from 500M to 7B, and up to 100B pretraining tokens finds that synthetic pre-pretraining (PPT) still improves token efficiency at scale — saving at least 21B PT tokens at the 3B scale.

Takeaway: PPT is a low-cost pretraining add-on, and future task design should target long-range retrieval rather than grammar.

Related event: Synthetic pre-pretraining scales to 7B models, saving 21B tokens(2 posts)→

Original post →

More from Research

Research channel →