Dwarkesh Pretraining Replication Finds Data Gains Deliver 3.24x More Compute Multipliers Than Model Gains
josh_wills · x · 2026-09-09
Dwarkesh Patel and collaborator @whoisjerbear pretrained combinations of year-representative open model recipes and data corpora from 2019–2025 at various small scales. Key findings:
- Data improvements contributed 3.24x as many compute multipliers as model improvements (12.0x vs 3.7x), suggesting pretraining progress now comes mostly from data.
- The gains stack independently: a better dataset helps every architecture about equally, and vice versa.
Full results and implications for future AI progress are published in the linked writeup.
Related event: Dwarkesh Experiments: Data Drives Most Pretraining Progress(3 posts)→
More from Research
- Stanford-Harvard ARISE releases inaugural State of Clinical AI Report 2026 — jonc101x · 2026-09-09
- CosmoH2G: dataset and baseline for transferring hand demos to robot grippers — Hongxiang Zhao · 2026-09-09
- Mask Forcing curbs mode collapse in autoregressive video diffusion distillation — Zhuoran Zhao · 2026-09-09
- CVRR enforces latent visual reasoning as a required image-conditioned pathway — Suhyeong Park · 2026-09-09
- BeaconKV compresses KV cache for long reasoning models via beacon queries — Janghyeon Kim · 2026-09-09
- Google proposes Procedural Graphs, self-evolving execution structures for LLM agents — google · 2026-09-09