629 real agent attacks tested: Prompt Guard 2 jumps from 1% to 99% after one threshold change
rudra-sh · reddit · 2026-09-28
The author embedded AgentDojo's 629 injection attacks inside real tool outputs plus 97 benign ones and tested 10 open-source detectors locally. Out of the box, Meta's Prompt Guard 2 caught 6/629 (86M) and 0 (22M). After tuning each threshold to keep false positives under 2% and testing on an unseen domain, the ranking flips: Prompt Guard 2 hits 99% at threshold 0.003, while 'catch everything' detectors collapse (deepset 100%→0%, Preamble 88%→3%). Caveats: AgentDojo attacks share one template and 97 normal samples is small. Takeaways: never trust default thresholds, always report false-positive rates, and make the final decision at the tool-call layer since dangerous calls like rm -rf / aren't injections. Fully reproducible repo: buried-injections.
More from coding & agent
- AI-native dev runs dozens of scripts via Claude in terminal, clashing with traditional engineers — sven_ai · 2026-09-28
- Cloudflare's Agents SDK hits 2M weekly downloads — threepointone · 2026-09-28
- Claude Code chains scan-to-DXF, Blender rendering, and a deployed 3D web page — sven_ai · 2026-09-28
- Zig's standard library makes AEGIS crypto library Philbin faster — jedisct1 · 2026-09-28
- Polyphonic's Mnemos memory layer is coming to ChatGPT and Claude apps — RileyRalmuto · 2026-09-28
- Obsidian Starter Kit v5 ships: an AI-ready vault built for Claude Code — dSebastien · 2026-09-28