Pretraining gains come mostly from data: experiments show 12x vs 3.7x compute multipliers
eliebakouch · x · 2026-09-09
Dwarkesh Patel and a collaborator pretrained combinations of year-representative open model recipes and data corpora from 2019 to 2025 at various small scales to quantify where pretraining progress comes from.
- Data improvements delivered 3.24x as many compute multipliers as model improvements (12.0x vs 3.7x).
- The gains stack independently: a better dataset helps every architecture roughly equally, and vice versa.
- Full results and implications for the future of AI progress are published in a writeup.
The finding pushes back on the narrative that architecture innovation is the main engine of pretraining progress.
Related event: Dwarkesh Experiments: Data Drives Most Pretraining Progress(3 posts)→
More from Models
- Gary Marcus amplifies question: what exactly does OpenAI commit to when you toggle this setting off? — GaryMarcus · 2026-09-09
- Meta's Muse usage blows past projections: users consuming 10x more than test cohorts — alexandr_wang · 2026-09-09
- Abacus.AI teases near-free LLM targeting long-running personal agentic loops, launching Thursday — bindureddy · 2026-09-09
- Artificial Analysis launches Model Release pages comparing every effort level of frontier models — ArtificialAnlys · 2026-09-09
- GLM 5.3 Flash goes live on W&B serverless inference: 1M context, vision, $0.50/M output — wandb · 2026-09-09
- DeepSeek V4 Flash listed at $0.05/$0.16 per 1M off-peak on OpenRouter, 4x below official pricing — michaelsoft__binbows · 2026-09-09