Defending Finetuning Poisoning via Trusted LoRA Subspaces

Bright_Warning_8406 · reddit · 2026-07-08

The author proposes a novel defense against finetuning poisoning: constraining finetuning within a subspace learned from trusted LoRA adapters. This renders certain malicious update directions geometrically unreachable while preserving useful adaptation capabilities. Tested on 196 public LoRA adapters (including specially designed adaptive attacks), the attack success rate dropped significantly without compromising the adapter's task coverage. Both the paper and code are publicly available.

Original post →

More from Safety

Safety channel →