Poison Set Selection Can Skyrocket LLM Backdoor Attack Success
Anthropic research shows that learning to select poison samples, rather than sampling them randomly, can raise LLM backdoor attack success rates from around 3% to 80% on LLaMA-3-8B, suggesting current evaluations underestimate model vulnerability.
2026-09-15 ~ 2026-09-15 · 2 related posts
- Poison set choice swings LLM backdoor attack success from 3% to 80%, SAILS paper shows — chaumian · 2026-09-15
- Anthropic research: learned poison selection boosts LLM backdoor attack success — Anthropic · 2026-09-15