BeyondWeb: Lessons from Scaling Synthetic Data for Trillion-scale Pretraining
DatologyAI, :, Pratyush Maini, Vineeth Dorna, Parth Doshi, Aldo Carranza, Fan Pan, Jack Urbanek, Paul Burstein, Alex Fang, Alvin Deng, Amro Abbas, Brett Larsen, Cody Blakeney, Charvi Bannur, Christina Baek, Darren Teh, David Schwab, Haakon Mongstad, Haoli Yin, Josh Wills, Kaleigh Mentzer, Luke Merrick, Ricardo Monti, Rishabh Adiga, Siddharth Joshi, Spandan Das, Zhengping Wang, Bogdan Gaza, Ari Morcos, Matthew Leavitt
cs.LG, cs.CL
2025-08-15
BeyondWeb rephrases high-quality web text into diverse formats with small LLMs: +5.1pp over Cosmopedia on 14 benchmarks, reaching web-data accuracy 7.7x faster.
Pretraining ran into the data wall: the supply of high-quality, information-dense human text cannot keep up with trillion-token budgets, and repeating data only leads to overfitting. Synthetic data is the accepted answer, and every 2025-era frontier model reports heavy use of it. What has been missing is a controlled experimental account of why it works and which variables actually matter.
Two paradigms compete. Generator-driven approaches (Tiny Stories, Phi, Cosmopedia) prompt a large model to write educational content from parametric knowledge; they are expensive and inherit the generator's coverage gaps. Source rephrasing (WRAP, Nemotron-CC) uses smaller models to rewrite existing web documents into question-answer pairs and other task-aligned formats, and has become the industry default. DatologyAI puts both through a shared experimental setup, ablates seven variables one at a time, and assembles the findings into its own pipeline, BeyondWeb.
BeyondWeb starts from a high-quality subset of DCLM selected with the company's own classifiers, and rephrases it with small LLMs using three strategy families: format transformation (web pages into Q&A pairs), style modification (a more pedagogical tone that matches conversational inference use), and content restructuring to raise per-token information density. The training mix is 60% random RedPajama plus 40% synthetic. Infrastructure moved from Slurm on AWS Hyperpod to Ray plus vLLM on Kubernetes to parallelize rephrasing across trillions of tokens.
Every design choice is backed by an ablation: seed quality beats novelty, style should match the deployment distribution, and generation strategies must stay diverse to survive long training runs. No single factor moves the needle much on its own; the five-point gains come from stacking them.
Main results use Llama-3 architectures at 1B, 3B and 8B, averaged over 14 benchmarks in 0-shot and 5-shot settings:
| Setup | Baseline | BeyondWeb | Gap |
| 1B, 1T tokens | RedPajama 50.7% | 57.4% | +6.7pp |
| 3B, 180B tokens | RedPajama 53.5% | 60.8% | +7.3pp |
| 8B, 180B tokens | RedPajama 56.6% | 63.7% | +7.1pp |
| 8B, 180B tokens | Nemotron-Synth 61.1% | 63.7% | +2.6pp |
| 8B, 180B tokens | Cosmopedia 58.6% | 63.7% | +5.1pp |
The 8B model reaches RedPajama's 180B-token accuracy in 23.2B tokens (7.7x speedup) and Nemotron-Synth's in 66.2B (2.7x). The 3B model on BeyondWeb scores 60.8%, above the 8B model trained on Cosmopedia (58.6%).
The ablations carry most of the information, all run at 1B with 20B tokens. A single "summarize the following text" prompt hits 46.7%, nearly matching Cosmopedia's carefully seeded educational generation (47.1%), which says most of the generator-driven benefit is plain information condensation; BeyondWeb sits at 50.4%. In the data-wall experiment, repeating 10B tokens twice scores 45.5% and naive LLM continuation 46.2%, exactly the full-data upper bound of 46.2%, while BeyondWeb beats that bound by 4.2pp. The wall is breakable, but naive continuation does not break it. Rephraser family barely matters: OLMo-2-7B, Llama-3.1-8B, Phi-4-14B and Mistral-7B all land within about one point, and OLMo, the weakest on benchmarks (59.6%), produces the best synthetic data (49.9%). Gains saturate past 3B rephrasers: 1B to 3B adds 1.5pp, 3B to 8B only 0.4pp. Conversational text is just 3.67% of the web mix; upsampling it to 50% adds only 0.9pp and flattens past 20%.
This is one of the few systematic ablation studies of synthetic pretraining data with production backing: BeyondWeb is part of the 7T-token dataset behind ArceeAI's AFM4.5B. The practical guidance is concrete. You do not need a frontier generator; a 3B open model rephrases well enough. Prompt diversity is worth more than any single clever format, especially over long horizons. And reusing high-quality data as seed beats importing low-quality novelty. At per-token training prices, a 7.7x cut in tokens-to-accuracy is a direct cost number.
The authors concede several points. Cosmopedia holds only 27B tokens and had to be repeated for longer runs, which they argue is an inherent weakness of generator-driven data but does handicap that baseline. In the continuation experiment, the generator has seen far more than the 20B tokens of the controlled setup and may inject parametric knowledge. The rephraser-size finding is flagged as specific to their setup.
Reading closely, the ablations mostly run at 1B with 20B tokens, so the jump to 8B and 180B rests on the headline runs rather than a full ablation grid. BeyondWeb's actual prompts and full recipe are not released, and infrastructure details are deferred to a follow-up, so reproduction is hard. DatologyAI sells a data curation platform, so the paper doubles as product literature. Benchmark contamination is never discussed.