ByteDance Studies Domain Data Repetition in LLM Pretraining
ByteDance-Seed · hf · 2026-08-17
ByteDance research investigates optimal repetition of high-quality domain data in LLM pretraining under proportional scaling of model size and tokens. Findings show optimal repetition increases mildly with scale and correlates more with domain validation loss than unique data volume.
More from Research
- Latent On-Policy Self-Distillation Improves Agent Performance — NationalUniversityofSingapore · 2026-08-17
- Test shows invisible Unicode chars can remove Claude text watermarks — Available-Deer1723 · 2026-08-17
- UMiami's ConlangCrafter AI invents complete languages from scratch — begusgasper · 2026-08-17
- Study: Minor architecture choices cripple long context capabilities — rohanpaul_ai · 2026-08-17
- Apple's New Benchmark Reveals LLMs Can't Do Math, Just Pattern Matching — anirbanbandyo · 2026-08-17
- Contactless Blood Pressure Monitoring Breakthrough: FMCW Radar with Momentum-Driven Calibration Achieves High Precision — maier_ak · 2026-08-17