Harvard & MIT Study: Data Reuse Leads to Diminishing Returns in LLM Training

burkov · x · 2026-08-13

LLM training is reaching a point where finding high-quality fresh text is harder than simply adding compute, while traditional 'scaling laws' largely assume fresh data can keep increasing with compute.

A recent preprint from Harvard and MIT examines a more realistic scenario: supplementing a fixed dataset with reused or paraphrased text. The authors introduce a metric to measure the value of these derived tokens compared to genuinely new ones.

Experiments show that the value of repeated data is not fixed: it yields diminishing returns as more is added. This provides a new perspective on how future training will be constrained by compute, fresh data, and the efficiency of data reuse.

Related event: Harvard & MIT Study: Data Duplication Wastes 33% of Compute in LLMs(3 posts)→

Original post →

More from Research

Research channel →