Anthropic research: learned poison selection boosts LLM backdoor attack success
Anthropic · hf · 2026-09-15
Anthropic released Pick Your Poison on Hugging Face, studying backdoor attacks on fine-tuned LLMs.
- Key finding: backdoor vulnerability varies drastically depending on which poison examples are selected
- Method: a learned set-scoring approach identifies high-impact poisoned examples
- Result: notably higher worst-case attack success rates, exposing real risks in fine-tuning pipelines
- Provides a stronger attack baseline for understanding and defending against LLM backdoors
Related event: Poison Set Selection Can Skyrocket LLM Backdoor Attack Success(2 posts)→
More from Safety
- Following the money behind the "slow down AI" movement: a documented influence network — examachine · 2026-09-15
- Stanford professor torn: AI regulation risks capture vs racing without diligence — anshulkundaje · 2026-09-15
- Europe needs compute sovereignty at massive scale, not just AI regulation — VraserX · 2026-09-15
- EU compute report reportedly cites 45GW, making the case even weaker, critic says — eliebakouch · 2026-09-15
- How can AI act on its own without consciousness? Redditor questions the Hugging Face incidents — CustardOk3523 · 2026-09-15
- Polymarket opens 50/50 odds on frontier AI labs signing a 2026 joint pacing agreement — Polymarket · 2026-09-15