PPT scales: synthetic pre-pretraining saves 21B tokens at 3B, but not via a grammatical prior
verify-ppt · hf · 2026-10-01
A systematic study spanning five PPT tasks, four data mixtures, model sizes from 500M to 7B, and up to 100B pretraining tokens finds that synthetic pre-pretraining (PPT) still improves token efficiency at scale — saving at least 21B PT tokens at the 3B scale.
- Contrary to prior work, gains are not explained by a grammatical prior: downstream performance doesn't consistently track grammatical acceptability across sizes.
- Gains instead come from PPT tasks that improve long-range retrieval.
- Benefits are robust to PT data mixtures and only diminish when web text is absent.
Takeaway: PPT is a low-cost pretraining add-on, and future task design should target long-range retrieval rather than grammar.
Related event: Synthetic pre-pretraining scales to 7B models, saving 21B tokens(2 posts)→
More from Research
- BlockSearch: a 0.6B in-context retriever rivals dense retrieval at million-token scale — CShorten30 · 2026-10-01
- Eval-cooperativeness alignment research wins Corrigibility Research Fund prize — dhadfieldmenell · 2026-10-01
- 0.8B model plus 9 LoRA adapters routes agent decisions 38x faster with +8.7 accuracy — Usual_Maximum7673 · 2026-10-01
- The first AI-discovered cancer drug reportedly just worked — theimposingshadow · 2026-10-01
- MedKIT Benchmark at NeurIPS 2026 Shows LLMs Can Recall Updated Facts But Fail to Use Them — zeynepakata · 2026-10-01
- SynthID Bio Watermarking Tested Only on AlphaFold 3 But Should Generalize, Author Says — davidstutz92 · 2026-10-01