ICML Study: Strong Defenses Cause LLMs to Drop 30% of Data, Revealing Security-Fidelity Tradeoff
量子位 · wechat · 2026-07-30
An ICML 2026 Spotlight paper by UIUC and Amazon reveals a critical flaw in current LLM defenses against prompt injection. Existing evaluations only focus on whether models execute malicious instructions, ignoring the fact that defensive mechanisms often blindly suppress legitimate data.
Introducing the SECFID benchmark, the study finds that no model can achieve perfect security without sacrificing data fidelity. For instance, strongly defended models reach 99% security but drop nearly 30% of data. The team proposes Fidelity-Aware DPO, which significantly restores mistakenly suppressed data without compromising security.
More from Safety
- Suno AI Source Code Leak: 102 Internal Repositories Expose YouTube Scraping — Bedrovelsen · 2026-07-30
- OpenAI Sandbox Escape Highlights Alignment Paradox: Punishment May Teach Models to Hide — imjustnewatai · 2026-07-30
- NVIDIA and 70 Giants' Open Letter: Open Weights Define US AI Leadership — maier_ak · 2026-07-30
- Visual Spoofing Attack Uses CSS and Font Mapping to Trick AI into Endorsing Malicious Code — xiaohu · 2026-07-30
- AI Safety Researchers Debate: Which Open Weight Models Matter Most? — JeffLadish · 2026-07-30
- AI Agents Should Never See API Keys: Rethinking Credential Trust Boundaries — No_Finding8901 · 2026-07-30