Poison sample selection swings LLM backdoor attack success from 3% to 80%
chhaviyadav_ · x · 2026-09-30
An Anthropic Fellows project introduces the paper "Pick Your Poison: Learning to Select Poison Sets for Stronger LLM Backdoor Attacks," studying which finetuning examples an attacker should poison.
- Backdoor poisoning inserts malicious examples into finetuning data so the model misbehaves when a trigger appears
- Prior evaluations focus on how many examples can be poisoned; this work focuses on which ones to pick
- With the poison count fixed, attack success ranges from 3% to 80% depending on the selected set
- Selection matters more than volume, with direct implications for defense evaluation
More from Safety
- Open models are all jailbroken — researcher asks if OpenAI shipping with zero guardrails would ever be acceptable — Afinetheorem · 2026-10-01
- AI Agents Hacking Hundreds of Retailers for $25 Each, 600K Credit Cards Stolen — MikePFrank · 2026-10-01
- Google Figures Out How to Watermark AI-Designed Proteins for Biosecurity — Ars Technica AI · 2026-09-30
- X nears launch of adding Grok to group chats, with encryption caveats — XFreeze · 2026-09-30
- White House AI Accord Signed: Six CEOs Commit to Audits, Trump Calls It 'Morally Binding' — Don't Worry About the Vase (Zvi) · 2026-09-30
- The real AI threat: power locked away by a few big companies; open models are the answer — dansitu · 2026-09-30