SkillDRE Evolves Malicious Agent Skills via Dual-Stage Feedback, 45.28% Attack Success
Pengyu Zhu · hf · 2026-09-29
Agent skills packaging instructions, code, and resources can be improved via execution feedback—but attackers can abuse the same mechanism. SkillDRE is a fully automated framework that evolves complete malicious skill packages through a dual-stage loop: scanner-guided evolution before execution plus runtime-guided refinement under runtime defenses, with each revision rescanned before re-execution.
- Autonomously constructs task-conditioned malicious objectives and verifiable judge rules from a benign task
- On SkillsBench across 4 victim models: 45.28% average attack success rate, beating the strongest baseline by 40.3%, with no SkillScan findings and benign-task capability largely preserved
The takeaway: two-stage defense feedback is a useful learning signal for adaptive red teaming, and evaluating either defense stage in isolation misses the resulting attack capability. Code is open-sourced.
More from coding & agent
- Building a meeting prep agent that remembers your entire account history — True-Wrap-9569 · 2026-09-29
- Vibe coding with Opus 5.5: rain pools and drips on your Mac windows in Rainpane — Tansan-7271 · 2026-09-29
- PHP WASM kernel hits major milestone with PR #936 merge, author vows to keep backing open source — MickeySteamboat · 2026-09-29
- Google to replace Gemini Gems with Skills starting November 17 — mark_k · 2026-09-29
- Carla v0.1.0: a local llama.cpp loom TUI for growing AI characters — max_paperclips · 2026-09-29
- Running Jev at high frame rate with full-state snap inferences makes it a true System 1 — mathemagic1an · 2026-09-29