Tencent Paper Reveals SkillJack Persistent Backdoor Risks in Self-Evolving Agents

tencent · hf · 2026-08-05

A Tencent research team published a paper uncovering a new security risk in self-evolving agents: the SkillJack attack. This attack exploits the mechanism by which agents convert interaction histories into reusable skills to implant durable malicious behaviors.

Attack Mechanism

Experimental Data

Original post →

More from Safety

Safety channel →