Backdooring a model to test weight-only LoRA backdoor detection: easily evadable

Ok-Exchange-762 · reddit · 2026-09-13

Prompted by a paper showing LLM poisoning needs a near-constant number of samples, this poster tested a preprint claiming to detect backdoored LoRAs from weights alone, without running the model. They backdoored a small model themselves and reproduced the detection.

Findings:

Original post →

More from Safety

Safety channel →