Synthetic pre-pretraining scales to 7B models, saving 21B tokens
The largest-scale pre-pretraining study to date shows PPT remains effective up to 7B parameters, saving about 21B tokens, but works via long-range retrieval rather than syntactic priors.
2026-10-01 ~ 2026-10-01 · 2 related posts
- Synthetic Pre-pretraining Survives Scaling to 7B, but Gains Come From Long-range Retrieval, Not Grammar — nikaletras · 2026-10-01
- PPT scales: synthetic pre-pretraining saves 21B tokens at 3B, but not via a grammatical prior — verify-ppt · 2026-10-01