Commenter warns LLM outputs poison training data, even modest distillation may be irreversible

teortaxesTex · x · 2026-09-21

A commenter raises a notable concern: model outputs (e.g. Claude's text) are forming priors for other LLMs, and even modest distillation may irreversibly poison the dataset. They argue this scaffolding for efficient RL will be amplified and entrenched by further legitimate internal RL, and it's unclear how deep the effect goes — adding a tongue-in-cheek note that Anthropic staff may already all talk the same way.

Original post →

More from AGI Musings

AGI Musings channel →