ICML Paper Reveals Data Duplication Wastes 33% of Compute

A recent ICML paper reveals that internal data repetition in pretraining severely degrades model performance and wastes approximately 33% of compute. The study highlights that even after rigorous deduplication, semantic repetition persists, with larger models being disproportionately more vulnerable to this hidden cost.

2026-08-12 ~ 2026-08-12 · 2 related posts