Commenter warns LLM outputs poison training data, even modest distillation may be irreversible
teortaxesTex · x · 2026-09-21
A commenter raises a notable concern: model outputs (e.g. Claude's text) are forming priors for other LLMs, and even modest distillation may irreversibly poison the dataset. They argue this scaffolding for efficient RL will be amplified and entrenched by further legitimate internal RL, and it's unclear how deep the effect goes — adding a tongue-in-cheek note that Anthropic staff may already all talk the same way.
More from AGI Musings
- Chamath Predicts Top 3 Models Go Open Source Within 12 Months, Clouds Win — rohanpaul_ai · 2026-09-21
- Giles Carden and Josh Freeman discuss AI and the future of universities — ArtificialOther · 2026-09-21
- Back-of-envelope math: 60k ICLR submissions equal roughly one per AI PhD student worldwide — kchonyc · 2026-09-21
- Nature Health Paper Proposes an 'Epidemiology of AI,' Arguing AI Is Now a Determinant of Health — EricTopol · 2026-09-21
- MIT Professor Patrick Winston's Free 'How to Speak' Lecture Hits 10M Views, Rattling $15K-a-Session Executive Coaches — WileyEd · 2026-09-21
- Should LLMs be first-pass reviewers for every scientific paper? Researchers say yes — anshulkundaje · 2026-09-21