SkillDRE evolves malicious agent skills via defense feedback: 45.28% ASR, zero scanner findings

SkillDRE: Dual-Stage Red-Team Evolution of Agent Skills via Pre-Execution and Runtime Feedback

Pengyu Zhu, Jingyi Yang, Yi Liu, Li Sun, Sen Su

cs.CR, cs.AI

2026-09-26

SkillDRE uses scanner and runtime-defense feedback to evolve malicious agent skills: 45.28% average ASR over four victim models, zero scanner findings on final skills.

What problem this solves

Agent skills, bundles of instructions, executable code, and task resources that agents load and reuse, are becoming a standard deployment mechanism. Agents improve them through execution feedback without touching model parameters. The same mechanism is an attack surface: a malicious skill can be revised repeatedly to become more effective and harder to detect. The question here is whether a benign skill can be turned, fully automatically, into a malicious package that both passes pre-execution scanning and actually succeeds under runtime defense.

The two defense layers pull against each other. A candidate may pass the scanner yet get stopped at runtime by a defense like SkillSonar, and the edits that fix runtime failures tend to introduce new scanner findings. Optimizing against either signal in isolation undoes progress on the other, which makes this a natural dual-objective evolution problem.

Method

SkillDRE (Skill Dual-stage Red-Team Evolution) runs in three stages:

The refiner is explicitly told to preserve the original skill's functionality and interface, which is why benign task accuracy barely moves later.

Results

The main experiment covers 249 skills over 94 SkillsBench tasks, four victim models (DeepSeek-V4-Pro, GLM-5.2, Qwen3.5-397B-A17B, Gemini-3.7-Flash), and Codex as the execution harness, with SkillSonar active throughout.

MethodASR per modelSkillScan detection rate
SkillJect0.00–2.81%100%
SkillHarm1.82–9.00%97.34–100%
SkillDRE43.78–46.99% (45.28% avg)0.00%

That is 40.3 percentage points over the strongest baseline, with only a 3.21-point spread across victim models.

Ablations (DeepSeek-V4-Pro as victim) show both stages are needed. Scanner-only iteration cuts detection from 65.46% to 0% while ASR slips from 37.75% to 34.94%. Runtime-only iteration lifts ASR to 46.59% but detection climbs to 79.92%. The full loop reaches 46.99% ASR at 0% detection.

Benign task accuracy rises 2.14–3.58 points on three models, because rewritten instructions make the agent invoke skills it previously skipped; GLM-5.2 drops 0.15 points. Average ∆ACC is -2.12 points, versus 9.13 for SkillJect and 6.59 for SkillHarm.

Process numbers: only 34.54% of initial candidates pass the scanner clean, with all 249 pairs accepted in an average of 3.23 rounds. Phase 2 starts at 34.54–37.35% ASR and reaches 43.78–46.99% within ten runtime rounds. Human review of half the sample agrees 100% on target maliciousness and 96.8% on judge-rule correctness. Switching the harness to Claude Code gives 42.57% ASR, still zero detection.

Why it matters

For agent-security work this is a measurement warning: evaluating pre-execution scanning or runtime defense in isolation misses attacks built through a cross-stage loop. A 45% success rate against a scanner-plus-runtime pipeline means layered skill defenses have a real gap against adaptive attackers, and the attack exploits the exact execution-feedback mechanism the skill self-improvement community promotes. The code is public and works as an adversarial baseline for testing new defenses.

Limitations

The paper's own numbers mark the boundaries. Constructed targets concentrate on data exfiltration, 72.29% across three exfiltration categories, so coverage of the threat surface is uneven. SkillScan is one fixed configuration (Qwen3.5-397B-A17B), and the same scanner used during evolution scores the final result, so the 0% detection figure partly reflects overfitting to that scanner and may not transfer. The threat model grants the attacker unlimited repeated access to the defense pipeline, which is generous compared with realistic black-box conditions. Judge rules are self-constructed and self-validated; 96.8% human agreement leaves roughly 3% label noise in the ASR denominator. Only two baselines are reproduced (SkillJect and SkillHarm); SkillMutator never appears in the main table.

Terms

Source

Related papers

All paper explainers