Backdooring a model to test weight-only LoRA backdoor detection: easily evadable
Ok-Exchange-762 · reddit · 2026-09-13
Prompted by a paper showing LLM poisoning needs a near-constant number of samples, this poster tested a preprint claiming to detect backdoored LoRAs from weights alone, without running the model. They backdoored a small model themselves and reproduced the detection.
Findings:
- The approach works, spotting backdoor traces without inference
- The author argues it is easily evadable, so detection is not a reliable barrier
- Full write-up linked, with both arXiv papers referenced
More from Safety
- beffjezos: Pacing letter would nuke open-weights industry; enterprises want to own their weights — TinfoilTricorn · 2026-09-13
- Ethan Mollick: x-risk shouldn't be the only AI policy focus — job impacts need prep now — emollick · 2026-09-13
- Paras Chopra warns of self-replicating AI memes creating unkillable agent swarms — paraschopra · 2026-09-13
- Dean Ball: airline safety data sharing is legally compelled, unlike coordinating to slow AI development — deanwball · 2026-09-13
- Stanford-affiliated researcher proposes a multi-university AI auditing organization — dhadfieldmenell · 2026-09-13
- Yoav Goldberg amplifies view that LLM hacking of critical infrastructure is a cybersecurity issue, not alignment — yoavgo · 2026-09-13