Prompt injection detection benchmarks have badly mislabeled data, researcher finds

evilsocket · x · 2026-10-10

Security researcher evilsocket examined one of the main datasets used to evaluate prompt injection detection and found samples labeled as injections that clearly aren't — arguing that many in the AI era have forgotten basic data analysis and QA, and that such benchmarks and product claims built on them deserve little trust.

Related event: Researcher Finds Major Labeling Errors in Prompt Injection Benchmarks(2 posts)→

Original post →

More from Safety

Safety channel →