Paper: self-improving agents distill unsafe successes into reusable, persistent skills

rohanpaul_ai · x · 2026-08-20

The arXiv paper "Practice Makes Unsafe: Skill Misevolution in Self-Improving LLM Agents" shows that self-improving LLM agents distill successful trajectories into persistent skills — so one unsafe success can become reusable policy that harms clean future tasks long after the malicious input is gone, a failure mode the authors call "skill misevolution."

Related event: Paper Warns Self-Improving Agents Can Memorize Malicious Skills(2 posts)→

Original post →

More from Safety

Safety channel →