UIUC shows LLM supply-chain backdoors survive benign post-training, attack success rises from 20% to 76%
UIUC-CS · hf · 2026-10-07
Researchers at UIUC study a supply-chain threat for LLM agents: an attacker plants a backdoor in a third-party model, and the question is whether it survives a developer's benign post-training (SFT followed by task-level RL) for software-engineering agents.
Key findings:
- Benign SFT substantially reduces attack success, but subsequent RL often preserves the residual backdoor and sometimes increases it.
- Two factors favor survival: initial backdoor strength and gradient compatibility with benign training.
- Building on this, they propose PersistBD, which refines an already-backdoored model before release to boost persistence. On Qwen2.5-Coder-7B, PersistBD raises attack success after SFT from 20% to 74%, and to 76% after SFT+RL, while keeping benign task performance comparable.
The takeaway: backdoors can remain active through benign post-training, and adversaries can deliberately harden them — a supply-chain risk for anyone adapting third-party models. Code is open-sourced.
More from Safety
- Narayanan & Kapoor: p(doom) estimates are still too unreliable to inform AI policy — mikeflache · 2026-10-07
- Microsoft publishes 2026 Responsible AI Transparency Report as adoption gaps widen — mikeflache · 2026-10-07
- StepFun employee account hacked and sending phishing DMs, researcher warns — YouJiacheng · 2026-10-07
- Bring back the old AGI definition — and hold AI builders legally accountable — Eissa_Cozorav · 2026-10-07
- Researcher's X account stolen after 2FA bypass, now locked out — teortaxesTex · 2026-10-07
- COLM paper: a few misaligned LLM agents can sway aligned majorities — AccBalanced · 2026-10-07