SEAD: SAGE defender cuts tool-agent attack success from 48% to 4% against DART attacks
Xinjie Shen · hf · 2026-09-30
SEAD formalizes attack and defense for tool-using agents as partially observed state control: earlier actions can alter files, permissions, or database state, making later routine-looking actions harmful in ways the visible interaction doesn't reveal.
- DART (attacker) decomposes harmful goals into locally plausible steps and uses real tool-execution feedback to guide trajectory search.
- SAGE (defender) investigates relevant state via read-only queries before allowing or blocking each action, including ones proposed after a block.
- Comes with an environment-verifiable dataset with controlled initial states, replayable tool environments, and executable task checks.
- Results: DART improves semantic attack success by 18.8–35.9 points over baselines across four target models; SAGE preserves 95.79% of benign trajectories, intercepts 92.73% of harmful ones, and cuts DART's executable attack success from 48.0% to 4.0%, generalizing across four attack methods and out-of-domain environments.
Code and data: https://github.com/EverywhereSafety/SEAD
More from coding & agent
- Free bookmarklet rebuilds full Google-rendered page in Search Console — gaganghotra_ · 2026-09-30
- Updated Dharmamitra Emacs package now available on MELPA — SebastianNehrd2 · 2026-09-30
- mitsuhiko mocks the reality of "open standards": great in theory, messy in practice — mitsuhiko · 2026-09-30
- DHH: AI-generated code can be hilariously hideous—it's just a prompt compilation target — mitsuhiko · 2026-09-30
- The gap between tutorial toy code and production AI systems is 'genuinely depressing' — Top-Philosopher-5411 · 2026-09-30
- Cloudflare's MCP redesign cuts tool context from 244K to 1.1K via catalog + executor — PuzzledFarmer4554 · 2026-09-30