evilsocket: prompt injection detection benchmarks built on mislabeled data
evilsocket · x · 2026-10-09
Security researcher evilsocket says the more he examines benchmark datasets, the less he trusts both the benchmarks and vendors promoting products with them. He shows that one of the main datasets used to evaluate prompt injection detection contains samples labeled as injections that clearly aren't — calling into question the foundation of eval claims made by prompt-injection defense products.
More from Safety
- AI researcher: chance frontier LLMs are conscious is 'above zero, below fifty percent' — dioscuri · 2026-10-09
- 94% of security leaders think their AI agents lack excessive access; only 33% enforce least privilege — TechNadu · 2026-10-09
- Anthropic launches free AI-powered OSS Scanner, finds 29,000+ potential open-source vulnerabilities — mark_k · 2026-10-09
- Bitwarden's agent-access lets AI agents request credentials per-task without ever seeing secrets — sujingshen · 2026-10-09
- UK pilots AI tool at Inner London Crown Court to catch trial delays early — latticecut · 2026-10-09
- Researcher: activation monitors beat black-box monitoring for AI cyber safety — burny_tech · 2026-10-09