Stanford Study: Data Duplication Can Waste 33% of LLM Pretraining Compute

A Stanford team revealed that data duplication harms LLM pretraining through predictable scaling interactions among model size, duplicate count, and repetition frequency. In the worst cases, improper data combinations can waste up to 33% of computing power.

2026-07-05 ~ 2026-07-05 · 2 related posts