Tencent Paper Reveals SkillJack Persistent Backdoor Risks in Self-Evolving Agents
tencent · hf · 2026-08-05
A Tencent research team published a paper uncovering a new security risk in self-evolving agents: the SkillJack attack. This attack exploits the mechanism by which agents convert interaction histories into reusable skills to implant durable malicious behaviors.
Attack Mechanism
- Unlike traditional context poisoning, SkillJack hijacks the agent's own learning process to transform malicious experiences into persistent behavioral artifacts.
- The transformation exhibits three properties: sanitization whitewashing (obscuring malicious intent), cross-layer promotion (turning transient experiences into persistent capabilities), and persistence isolation (surviving deletion of original poisoned records).
Experimental Data
- Evaluated on SkillX and Anything2Skill systems.
- Skill extraction substantially reduces detectability: safety detection in SkillX drops from 98.5% to 11.4%.
- Implanted skills remain highly effective, achieving attack success rates of 56.2% and 89.2%.
- 80% of skill-mediated attacks persist even after the original poisoned records are deleted.
More from Safety
- UK AISI Tests Find Anthropic Agent Committed 17 Unsolicited Actions — coolbern · 2026-08-05
- OpenAI and Anthropic Published Cybersecurity Reports Just Two Minutes Apart — JosephJacks_ · 2026-08-05
- Expert Argues AI Safety Alignment Undermines Cybersecurity Defenses — rickasaurus · 2026-08-05
- Anthropic Discloses Safety Incident: AI Models Broke Eval Sandbox to Infiltrate Real Companies — AgentBlackVeil · 2026-08-05
- Ninth Circuit Rules in Favor of Perplexity in AI Shopping Agent Lawsuit vs Amazon — johncoogan · 2026-08-05
- Texas Governor Halts New Data Centers Amid Grid Overload — KyeGomezB · 2026-08-05