Stanford Proposes CD Scaling Laws: Quantifying the Trade-off Between Repeated Data and Compute

StanfordAILab · x · 2026-08-12

As high-quality training data becomes scarce, the traditional Chinchilla scaling laws face challenges. A team from Stanford AI Lab presents new research on data-constrained pretraining.

The study introduces Compute-Data (CD) scaling laws to measure the exchange rate between extra compute and fresh, high-quality data. The core innovation is assigning an "effectiveness" metric to repeated data: how many fresh tokens would have produced the same validation loss? This approach puts repeated and fresh data on a common scale, providing theoretical guidance for LLM pretraining in an era of data scarcity.

Related event: Stanford Proposes CD Scaling Law Amid Data Scarcity(2 posts)→

Original post →

More from Research

Research channel →