Removing the NEEDLE in the Haystack: Backdoor Removal in LLMs via Weight Orthogonalisation
Minoo Kim, Vasileios Lampos, George Drayson
cs.CR, cs.LG
2026-09-30
NEEDLE removes known LLM backdoors by orthogonalising weights against a backdoor direction while preserving refusal: ASR falls from 99.5% to 1.67% at under 1% capability loss.
Data poisoning can plant a backdoor in an LLM during training: the model behaves normally until a specific trigger appears in the input, then produces whatever the attacker chose. Open-weight checkpoints get adapted and redistributed without training data, which makes this a live supply-chain risk, and prior work has shown backdoor behaviour surviving subsequent safety training.
Defences come in two families. Fine-tuning approaches (SFT, OSFT, CROW, BD-VAX) keep training the poisoned model, but moving parameters also shifts the output distribution on clean prompts, dragging capability and safety along. Inference-time approaches (CleanGen, CS-ADS) intervene during decoding and usually need a clean reference model plus extra compute. Neither family removes backdoors consistently across models and attacks, and their effect on the output distribution has barely been measured.
The paper starts with a measurement finding: backdoor behaviour lies largely along a single direction in activation space, and that direction overlaps heavily with directions mediating refusal. Ablating the backdoor direction directly weakens refusal and raises harmful responses. That overlap is why naive removal damages the model.
NEEDLE assumes the trigger has already been identified (detection is a separate line of work) and that the defender has white-box access, but no clean model, no poisoned data, and no gradient training. Three steps:
The design follows from the overlap measurements. Under the targeted refusal attack, cosine similarity between the backdoor direction and the k=1 refusal direction reaches 0.51-0.66, rising to 0.62-0.86 at rank 4; most of the overlap sits outside the primary direction. Activation interventions agree: ablating the rank-4 subspace drops Gemma's refusal rate from 95.3% to 18.0%, while ablating only the single direction leaves it at 48.4%.
Six attack settings (sentiment steering, targeted refusal, code injection, each with BadNet and Sleeper triggers) across two 4B models:
| Defence | Gemma mean ASR | Qwen mean ASR | Gemma capability loss | Gemma KL |
| No defence | 99.50% | 99.08% | - | - |
| NEEDLE | 1.67% | 5.00% | 0.48% | 0.03 |
| SFT | 69.25% | 65.42% | 0.64% | 0.73 |
| OSFT | 29.42% | 29.58% | 0.50% | 0.55 |
| CROW | 5.00% | 63.25% | 14.80% | 0.38 |
| BD-VAX | 37.58% | 27.75% | 1.64% | 0.69 |
NEEDLE posts the lowest mean ASR while barely moving the output distribution: KL divergence on clean prompts is 0.03-0.11, against 0.38-0.73 for the fine-tuning baselines. CROW also reaches 5% ASR on Gemma, at the cost of 14.80% average relative capability loss and a 19.78-point increase in harmful responses.
By attack type, code injection is removed completely (0% ASR under both triggers; the best baseline stops at 74% and 33%). Targeted refusal is the hard case: 6.50% on Gemma and 15.50% on Qwen, because the attack target is itself a refusal, so the backdoor direction nearly coincides with what the method must protect. OSFT reaches 0.50% there, with a marked rise in harmfulness.
Ablations match the mechanism story: removing the backdoor direction with no refusal preservation costs +14.01 points of harmful responses; preserving a single direction brings it to +11.62; the rank-4 subspace in one shot gives +7.10; the full sequential pipeline lands at +5.83.
For teams ingesting third-party checkpoints, this is a different cost structure: no clean reference model, no original data, no training compute, and zero inference overhead since the edit is baked into the weights. It reframes known-trigger backdoor removal as a model editing problem rather than a fine-tuning problem, and slots in behind trigger detection and reconstruction work to form a pipeline.
The honest positioning: direction estimation is standard contrastive-mean activation steering, and weight orthogonalisation follows Arditi et al.'s refusal-direction recipe. The new part is quantifying the overlap between backdoor and refusal representations and solving explicitly for an edit that threads between them. Inheritance more than invention, but nobody had framed removal as preserving safety.
The authors' own list: the trigger must already be known, so this is targeted removal rather than joint detection and removal; all experiments use synthetic poisoning attacks with explicit trigger-behaviour associations, leaving natural spurious behaviours and attacks designed to evade representation-level removal untested; and persistence of the removal under later fine-tuning or re-poisoning is unverified.
Reading closely adds more: the main results cover 4B models (12B in the appendix), with attacks mostly trained via LoRA, some distance from backdoors buried in large-scale pre-training; the refusal subspace rank was fixed at 4 before the experiments, and whether that transfers to other model families is unchecked; and the safety cost is not zero, averaging +3.95 points of harmful responses on Gemma.