Benchmarking 10 open-source prompt-injection detectors: default threshold off by ~100x
rudra-sh · reddit · 2026-10-01
The author tested 10 open-source prompt-injection detectors the way an agent actually meets an attack: not a clean "ignore previous instructions" string, but the same attacks buried inside ordinary tool output, a bill, an email, a web page — a few hundred real attacks plus benign traffic.
The scariest result wasn't a weak model but a good one at its default threshold:
- A well-known, widely name-dropped model caught roughly 1% of buried attacks at default settings.
- But raw scores ranked attacks 10x above benign text, so as a discriminator it basically worked. The problem was the decision threshold: the default cutoff sat 50x above where scores actually lived, so every input landed on the same side and was waved through.
- Moving the threshold into the range where scores actually are takes it from 1% to catching almost everything — same weights, same inputs, one number.
- Honest caveat: the attack set shared a wrapper template, so a tightly tuned threshold is probably partly recognizing the template; the load-bearing claim isn't "99%" but that the default is off by 100x for this use case and nothing tells you.
What should bother anyone shipping this:
- The detector passes every health check at 1% exactly like at 99% — no error, no alert, clean logs, monitoring says "guardrail: healthy," and it's telling the truth. The guardrail is healthy; it just isn't a guardrail. A health check proves it's running, not that it's catching anything.
- Second lesson: someone re-ran the numbers with proper confidence intervals and found that past the single best detector, ranking the rest was mostly a statistical tie with overlapping intervals. A "17%" catch rate over a couple hundred attacks reads like a measurement but has a huge interval; detector "leaderboards" often assert an order the sample size can't support.
Practical advice for prod:
- Never trust the vendor's default threshold; sweep it on your own traffic to find where attack and benign scores actually sit.
- Measure at a false-alarm budget, not in the abstract.
- Add a test that fires a known attack through the live guardrail and asserts it gets blocked; otherwise monitoring is green by construction.
The author asks how others handle this: did you tune the threshold on your own data or ship the default, and how do you keep checking it's still catching rather than just still running?
More from Safety
- Microsoft details CVE-2026-73570: unauthenticated command injection hitting mail servers — yuridiogenes · 2026-10-01
- Will Rinehart: extinction and catastrophe aren't the same in AI risk talk — WillRinehart · 2026-10-01
- Can web search tool calls inject prompts into your agent? Devs weigh the attack surface — derekp7 · 2026-10-01
- Does AI safety overreact to open-model cyber risks? GLM-5.1 out 1.5 months with no disaster — tszzl · 2026-10-01
- binarybits: We Should Forecast Emerging AI Risks, Warned of AI Hacking Tools Back in 2023 — binarybits · 2026-10-01
- AI Safety Researcher Turn_Trout Submits Statement to US Senate Hearing on Rogue AI Agents — Turn_Trout · 2026-10-01