Dwarkesh Experiments: Data Drives Most Pretraining Progress

Dwarkesh Patel and collaborators replayed 2019-2025 open-source model recipes and data at small scale, finding that data improvements contributed a compute multiplier about 3.24x that of model-side changes—roughly three-quarters of pretraining progress came from data (12x vs 3.7x).

2026-09-09 ~ 2026-09-09 · 3 related posts