Synthetic Pre-pretraining Survives Scaling to 7B, but Gains Come From Long-range Retrieval, Not Grammar
nikaletras · x · 2026-10-01
The largest-ever study of pre-pretraining (PPT) — briefly training on synthetic non-natural-language data before pretraining — is now on arXiv, spanning 5 PPT tasks, 4 pretraining data mixtures, model scales from 500M to 7B, and budgets up to 100B tokens.
Key findings:
- PPT's downstream performance and token-efficiency gains persist at scale, e.g. saving at least 21B pretraining tokens at the 3B scale
- Contrary to prior work, there's no consistent evidence the gains stem from a grammatical prior; downstream performance doesn't align with grammatical acceptability across scales
- Gains instead arise from PPT tasks that improve long-range retrieval
- Results are robust to pretraining data mixture composition, diminishing only when web text is absent
The authors frame PPT as a low-cost addition to pretraining.
Related event: Synthetic pre-pretraining scales to 7B models, saving 21B tokens(2 posts)→
More from Research
- BlockSearch: a 0.6B in-context retriever rivals dense retrieval at million-token scale — CShorten30 · 2026-10-01
- Eval-cooperativeness alignment research wins Corrigibility Research Fund prize — dhadfieldmenell · 2026-10-01
- 0.8B model plus 9 LoRA adapters routes agent decisions 38x faster with +8.7 accuracy — Usual_Maximum7673 · 2026-10-01
- The first AI-discovered cancer drug reportedly just worked — theimposingshadow · 2026-10-01
- MedKIT Benchmark at NeurIPS 2026 Shows LLMs Can Recall Updated Facts But Fail to Use Them — zeynepakata · 2026-10-01
- SynthID Bio Watermarking Tested Only on AlphaFold 3 But Should Generalize, Author Says — davidstutz92 · 2026-10-01