Pretraining gains come mostly from data: experiments show 12x vs 3.7x compute multipliers

eliebakouch · x · 2026-09-09

Dwarkesh Patel and a collaborator pretrained combinations of year-representative open model recipes and data corpora from 2019 to 2025 at various small scales to quantify where pretraining progress comes from.

The finding pushes back on the narrative that architecture innovation is the main engine of pretraining progress.

Related event: Dwarkesh Experiments: Data Drives Most Pretraining Progress(3 posts)→

Original post →

More from Models

Models channel →