629 real agent attacks tested: Prompt Guard 2 jumps from 1% to 99% after one threshold change

rudra-sh · reddit · 2026-09-28

The author embedded AgentDojo's 629 injection attacks inside real tool outputs plus 97 benign ones and tested 10 open-source detectors locally. Out of the box, Meta's Prompt Guard 2 caught 6/629 (86M) and 0 (22M). After tuning each threshold to keep false positives under 2% and testing on an unseen domain, the ranking flips: Prompt Guard 2 hits 99% at threshold 0.003, while 'catch everything' detectors collapse (deepset 100%→0%, Preamble 88%→3%). Caveats: AgentDojo attacks share one template and 97 normal samples is small. Takeaways: never trust default thresholds, always report false-positive rates, and make the final decision at the tool-call layer since dangerous calls like rm -rf / aren't injections. Fully reproducible repo: buried-injections.

Original post →

More from coding & agent

coding & agent channel →