Synthetic pre-pretraining scales to 7B models, saving 21B tokens

The largest-scale pre-pretraining study to date shows PPT remains effective up to 7B parameters, saving about 21B tokens, but works via long-range retrieval rather than syntactic priors.

2026-10-01 ~ 2026-10-01 · 2 related posts