UIUC shows LLM supply-chain backdoors survive benign post-training, attack success rises from 20% to 76%

UIUC-CS · hf · 2026-10-07

Researchers at UIUC study a supply-chain threat for LLM agents: an attacker plants a backdoor in a third-party model, and the question is whether it survives a developer's benign post-training (SFT followed by task-level RL) for software-engineering agents.

Key findings:

The takeaway: backdoors can remain active through benign post-training, and adversaries can deliberately harden them — a supply-chain risk for anyone adapting third-party models. Code is open-sourced.

Original post →

More from Safety

Safety channel →