New scaling law paper says repetition can beat paraphrasing for some pretraining regimes
burny_tech · x · 2026-07-29
A new paper proposes compute-data scaling laws to bridge the gap between compute-optimal and data-optimal pretraining.
- Classical scaling laws assume unlimited fresh data, but the paper argues pretraining is increasingly constrained by data availability.
- The authors introduce a token-effectiveness function η that measures how valuable derived tokens are relative to fresh ones.
- They fit the framework on Dolma-3 across model sizes from 14M to 600M parameters, studying two expansion strategies:
- multi-epoch repetition
- paraphrasing
- Key findings:
- token effectiveness is not constant; it depends on model size, tokens-per-parameter ratio, and derived-data volume;
- effectiveness saturates as the corpus expands;
- adding compute in place of data has diminishing returns as models and data get larger;
- classical compute-optimal allocation is suboptimal across most practical settings.
- The framework divides training into compute-bound, data-bound, and model-bound regimes and suggests repetition can beat paraphrasing in some model/corpus-size ranges.
More from Infra
- SK hynix reportedly signs long-term contracts with about 10 major customers — dejavucoder · 2026-07-29
- Shanghai Aishengna is said to be manufacturing DUV lithography tools — zephyr_z9 · 2026-07-29
- A vendor-agnostic Vulkan backend cuts edge inference latency from 30 ms to 3 ms — ppchaos · 2026-07-29
- pdf-mcp turns technical PDFs into structured text, images, and searchable context — tom_doerr · 2026-07-29
- Moonshot’s Kimi K3 is a 2.8T open-weight MoE model with 1M-token context — alex_verem · 2026-07-29
- SK hynix swings wildly as HBM profitability keeps rising — basedjensen · 2026-07-29