ICML Study: Strong Defenses Cause LLMs to Drop 30% of Data, Revealing Security-Fidelity Tradeoff

量子位 · wechat · 2026-07-30

An ICML 2026 Spotlight paper by UIUC and Amazon reveals a critical flaw in current LLM defenses against prompt injection. Existing evaluations only focus on whether models execute malicious instructions, ignoring the fact that defensive mechanisms often blindly suppress legitimate data.

Introducing the SECFID benchmark, the study finds that no model can achieve perfect security without sacrificing data fidelity. For instance, strongly defended models reach 99% security but drop nearly 30% of data. The team proposes Fidelity-Aware DPO, which significantly restores mistakenly suppressed data without compromising security.

Original post →

More from Safety

Safety channel →