MIT & Microsoft Paper Uncover Data Synergy in LLM Pretraining

LesterMackey · x · 2026-08-05

Machine learning progress is often attributed to scaling model size and dataset volume, but data composition is just as crucial. This paper from MIT and Microsoft Research formalizes and quantifies "data synergy" in language model pretraining.

The authors note that combining datasets from different domains yields nontrivial interactions—adding code improves math reasoning, while certain mixtures introduce interference. Leveraging observational variation across open-weight LLMs, they estimate direct domain-to-benchmark synergy and second-order domain-domain synergy. Their framework improves predictive accuracy over domain-agnostic scaling laws and successfully predicts performance rankings by training models on predicted optimal versus anti-optimal mixtures.

Original post →

More from Research

Research channel →