NEEDLE: training-free backdoor removal for LLMs drops code-injection attack success to 0%

locailabs · hf · 2026-10-02

Locai Labs released NEEDLE, a training-free method for targeted backdoor removal in LLMs. After identifying a trigger, it estimates a backdoor direction and refusal subspace from activation vectors, then applies sequential weight orthogonalisation to suppress the backdoor while preserving refusal-related representations — requiring neither a clean reference model nor poisoned training data. Across multiple model families and attack types, NEEDLE achieves the lowest mean attack success rate among evaluated defences, including 0% on code-injection attacks, with the lowest KL divergence and minimal capability/safety changes.

Original post →

More from Safety

Safety channel →