SkillDRE Evolves Malicious Agent Skills via Dual-Stage Feedback, 45.28% Attack Success

Pengyu Zhu · hf · 2026-09-29

Agent skills packaging instructions, code, and resources can be improved via execution feedback—but attackers can abuse the same mechanism. SkillDRE is a fully automated framework that evolves complete malicious skill packages through a dual-stage loop: scanner-guided evolution before execution plus runtime-guided refinement under runtime defenses, with each revision rescanned before re-execution.

The takeaway: two-stage defense feedback is a useful learning signal for adaptive red teaming, and evaluating either defense stage in isolation misses the resulting attack capability. Code is open-sourced.

Original post →

More from coding & agent

coding & agent channel →