evilsocket: prompt injection detection benchmarks built on mislabeled data

evilsocket · x · 2026-10-09

Security researcher evilsocket says the more he examines benchmark datasets, the less he trusts both the benchmarks and vendors promoting products with them. He shows that one of the main datasets used to evaluate prompt injection detection contains samples labeled as injections that clearly aren't — calling into question the foundation of eval claims made by prompt-injection defense products.

Original post →

More from Safety

Safety channel →