Poison set choice swings LLM backdoor attack success from 3% to 80%, SAILS paper shows
chaumian · x · 2026-09-15
A new paper shows backdoor attack evaluations drastically underestimate LLM vulnerability: across three LLaMA-3-8B settings, attack success ranges from 3% to 80% depending solely on which poison set is chosen. The authors formalize poison selection as oracle-budgeted set optimization and introduce SAILS, which learns a set scorer from a few hundred finetune-and-evaluate runs, ranks millions of candidate sets, and audits a shortlist. SAILS improves held-out attack success by 30 points over the strongest influence baselines, transfers to full-scale finetuning, and extends to code-generation, agentic, and API-only backdoors.
Related event: Poison Set Selection Can Skyrocket LLM Backdoor Attack Success(2 posts)→
More from Safety
- Cisco's VLoc Bench: even GPT-5.5 hits only 0.221 F1 at repo-scale vulnerability localization — aminkarbasi · 2026-09-15
- AistyMCP: open-source per-tool permissions for MCP servers, deny-by-default — iamjoehoward · 2026-09-15
- Anthropic co-founder backs mandatory AI 'kill switch'; commenter says make it law — srimisra · 2026-09-15
- Rejecting today's AI is demanding better tech, not resisting technology, argues philosopher — CarissaVeliz · 2026-09-15
- Lawmaker proposes one-month global AI safety stand-down to set red lines — DavidSKrueger · 2026-09-15
- User claims Claude Artifacts silently uploads drafts to cloud with toggle locked — maier_ak · 2026-09-15