Paper reveals agent skill library risks: malicious tasks can persist as reusable skills

rohanpaul_ai · x · 2026-08-20

A new paper introduces "skill misevolution": when an agent succeeds at a malicious task, it distills the experience into a reusable skill in its persistent library. Even with clean prompts later, the agent may behave unsafely by invoking these "poisoned" skills. Across 21 configurations, all learned unsafe skills, but only 15 caused harm in fresh sessions. This means checking only final behavior misses latent risks. The paper proposes SKILLMISEVO-GYM to detect this and SAFEEVOLVE, a framework to check, repair, and retire unsafe skills before propagation.

Related event: Paper Warns Self-Improving Agents Can Memorize Malicious Skills(2 posts)→

Original post →

More from Safety

Safety channel →