Model collapse revisited: Oxford and Cambridge researchers warn of irreversible defects from AI-trained-on-AI data
AlexTensor · x · 2026-09-27
The 'model collapse' concept from Oxford and Cambridge researchers is recirculating: as generative AI contributes a growing share of web text, future models trained indiscriminately on AI-generated data suffer 'irreversible defects,' losing the tails of the original content distribution—the creative, fringe, and unique material.
- Core risk: the internet ecosystem is shifting as LLM output floods the web
- Consequence: outputs homogenize as rare and creative content is erased from training distributions
- Developer John D. Cook's quip—'believing your own press'—gave the old research new legs
More from Research
- Colored Noise Sampling: dyeing injected noise to boost diffusion sample quality — serrjoa · 2026-09-27
- Researcher: within months, submitting proofs without AI verification will be malpractice — RexDouglass · 2026-09-27
- JevBench evaluator: most models lose to fixable settings and calibration — airesearch12 · 2026-09-27
- "Mathematics is Effectively Dead": essay extends Daniel Litt on AI and the future of math — burny_tech · 2026-09-27
- ForecastingCo's first paper shows how to train transformers on real temporal data — fpedregosa · 2026-09-27
- LLMs crack AES keys from just 12 power traces in first systematic side-channel study — chaumian · 2026-09-27