Poison set choice swings LLM backdoor attack success from 3% to 80%, SAILS paper shows

chaumian · x · 2026-09-15

A new paper shows backdoor attack evaluations drastically underestimate LLM vulnerability: across three LLaMA-3-8B settings, attack success ranges from 3% to 80% depending solely on which poison set is chosen. The authors formalize poison selection as oracle-budgeted set optimization and introduce SAILS, which learns a set scorer from a few hundred finetune-and-evaluate runs, ranks millions of candidate sets, and audits a shortlist. SAILS improves held-out attack success by 30 points over the strongest influence baselines, transfers to full-scale finetuning, and extends to code-generation, agentic, and API-only backdoors.

Related event: Poison Set Selection Can Skyrocket LLM Backdoor Attack Success(2 posts)→

Original post →

More from Safety

Safety channel →