New finding: AI-generated tokens raise loss once models pass Chinchilla-optimal data ratios
gleech · x · 2026-10-02
New research quantifies synthetic-data mixing: for data-starved models (under 10 human tokens per parameter), adding AI-generated tokens lowers loss on human text. But from Chinchilla-optimal (20 tokens/param) onward, adding AI tokens raises loss almost immediately, while human tokens keep lowering it — synthetic data's value depends heavily on how data-rich the model already is.
Related event: 800-Model Study: AI-Generated Text in Web Corpora Hurts Pretraining(5 posts)→
More from Research
- Meta paper: branching harness search lifts Olympiad math +34.8% over Meta-Harness — dair_ai · 2026-10-02
- ARPA-H plans to cut clinical trials from 10+ years to under 4 with AI — Afinetheorem · 2026-10-02
- Tacit-TTS: transcript-free voice cloning 10x faster than IndexTTS2 — Jian Chen · 2026-10-02
- ByteDance Seed's RWTD lifts one-step SANA Sprint GenEval from 0.73 to 0.80 — ByteDance-Seed · 2026-10-02
- Amazon's position-selective self-distillation trains LLM judges that beat RL by 2-9 points — amazon · 2026-10-02
- mhctools: one Python wrapper for a dozen MHC immunoinformatics predictors — iskander · 2026-10-02