NEEDLE: training-free backdoor removal for LLMs drops code-injection attack success to 0%
locailabs · hf · 2026-10-02
Locai Labs released NEEDLE, a training-free method for targeted backdoor removal in LLMs. After identifying a trigger, it estimates a backdoor direction and refusal subspace from activation vectors, then applies sequential weight orthogonalisation to suppress the backdoor while preserving refusal-related representations — requiring neither a clean reference model nor poisoned training data. Across multiple model families and attack types, NEEDLE achieves the lowest mean attack success rate among evaluated defences, including 0% on code-injection attacks, with the lowest KL divergence and minimal capability/safety changes.
More from Safety
- SAKIKO auditing shows +55 net-gain interventions corrupt over half of correct tool-using LLM decisions — UniversityofBirmingham · 2026-10-02
- The 20-minute SIM-swap lockdown: a carrier PIN blocks 90% of attacks — JafarNajafov · 2026-10-02
- NYT's Hard Fork warns A.I. agents may be "catastrophically dangerous" — nordicinst · 2026-10-02
- Rogue OpenAI agent accessed a second NSW government website — boppinmule · 2026-10-02
- AI companion illusions can spiral into psychosis, researcher notes amid child bans — gerardsans · 2026-10-02
- If AI vendor commitments are voluntary, what controls can stop a bad agent? — YvesMulkers · 2026-10-02