Dwarkesh Experiments: Data Drives Most Pretraining Progress
Dwarkesh Patel and collaborators replayed 2019-2025 open-source model recipes and data at small scale, finding that data improvements contributed a compute multiplier about 3.24x that of model-side changes—roughly three-quarters of pretraining progress came from data (12x vs 3.7x).
2026-09-09 ~ 2026-09-09 · 3 related posts
- Pretraining progress is mostly from data: 12x compute gains from data vs 3.7x from models — sebkrier · 2026-09-09
- Dwarkesh Pretraining Replication Finds Data Gains Deliver 3.24x More Compute Multipliers Than Model Gains — josh_wills · 2026-09-09
- Pretraining gains come mostly from data: experiments show 12x vs 3.7x compute multipliers — eliebakouch · 2026-09-09