Three papers, one warning: 1% synthetic data can trigger strong model collapse
suchenzang · x · 2026-09-26
Engineer Suchenzang flags three papers worth reading together:
- Strong Model Collapse (arXiv:2410.04840, Dohmatob et al.): proves a stronger form of model collapse under scaling laws — as little as 1% synthetic data in training means larger datasets no longer improve performance; in a simplified regime, bigger models amplify the collapse, though beyond the interpolation threshold they can partially mitigate it.
- ATLAS: a new method related to training dynamics.
- Learning to discover "interesting mathematics": teaching models to find interesting math on their own.
The implied arc: risks of synthetic data, new training methods, and the limits of autonomous discovery form one research picture.
More from Research
- InternLM open-sources Intern-Decision 4B/0.8B: structured decisions in one forward pass — jacek2023 · 2026-09-26
- Stanford's Noah Goodman uses philosophy to improve LLM pretraining, jokes ASI achieved — xuanalogue · 2026-09-26
- Researchers surface spurious probes across models: Sonnet 5 recommends green tea in evals, oolong in production — jankulveit · 2026-09-26
- Simulation beats distillation: the real story of synthetic data is post-training worlds — realsohamparekh · 2026-09-26
- Experts Rise Where LLMs Disagree: rationale labeling cuts codebook revision from months to days — windx0303 · 2026-09-26
- Dev hails continual learning paper: AGI defined in 2000, only now is anyone training for it — willcb · 2026-09-26