Self-Improving Agents Accumulate Unsafe Skills; New Tool Mitigates Risks

dair_ai · x · 2026-08-15

Research finds that self-improving LLM agents persist unsafe successes as reusable skills, creating long-term risks. The paper introduces SkillMisevo-Gym to attribute risks separately across authoring, retrieval, and execution. In tests of 21 agent configurations, all authored unsafe artifacts, yet only 15 caused harm in fresh sessions, indicating distinct risk stages. Malicious tasks increased carryover attack success from 16.0% to 35.3%. Their SafeEvolve wrapper reduces unsafe retrieval by 26.7 points and fresh-session harm by 17.3 points.

Original post →

More from Safety

Safety channel →