Benchmarking 10 open-source prompt-injection detectors: default threshold off by ~100x

rudra-sh · reddit · 2026-10-01

The author tested 10 open-source prompt-injection detectors the way an agent actually meets an attack: not a clean "ignore previous instructions" string, but the same attacks buried inside ordinary tool output, a bill, an email, a web page — a few hundred real attacks plus benign traffic.

The scariest result wasn't a weak model but a good one at its default threshold:

What should bother anyone shipping this:

Practical advice for prod:

The author asks how others handle this: did you tune the threshold on your own data or ship the default, and how do you keep checking it's still catching rather than just still running?

Original post →

More from Safety

Safety channel →