Pretraining progress is mostly from data: 12x compute gains from data vs 3.7x from models
sebkrier · x · 2026-09-09
- Dwarkesh Patel and Jerry Han pretrained combinations of year-representative open model recipes and data corpora from 2019-2025 at up to 1e19 FLOPs, evaluating end capabilities with OLMES rather than fixed-dataset loss.
- Data improvements delivered 12.0x in compute multipliers versus 3.7x from model improvements — data contributed 3.24x more to pretraining progress.
- Gains stack independently: better datasets help every architecture roughly equally and vice versa, with implications for frontier lab economics and the pace of future progress.
More from Research
- Terence Tao: The Question Actually Had a Small Finite Counterexample — tak3sh8 · 2026-09-09
- Uno paper: discrete diffusion drafting gives lossless LLM speedups without a draft model — rohanpaul_ai · 2026-09-09
- Uno speeds up Qwen3-8B 2.5x by using diffusion for parallel token drafting — rohanpaul_ai · 2026-09-09
- OpenAI claims AI found analytical proof of Navier-Stokes blowup, verified in Lean — Dr_Singularity · 2026-09-09
- LLMs' best robotics role: generating controllers to pretrain robot models — igilitschenski · 2026-09-09
- OpenAI Claims Navier-Stokes Millennium Prize Proof Produced by Agent Swarm on Next-Gen Model — daniel_mac8 · 2026-09-09