Defending Finetuning Poisoning via Trusted LoRA Subspaces
Bright_Warning_8406 · reddit · 2026-07-08
The author proposes a novel defense against finetuning poisoning: constraining finetuning within a subspace learned from trusted LoRA adapters. This renders certain malicious update directions geometrically unreachable while preserving useful adaptation capabilities. Tested on 196 public LoRA adapters (including specially designed adaptive attacks), the attack success rate dropped significantly without compromising the adapter's task coverage. Both the paper and code are publicly available.
More from Safety
- Open-source CLI audits AI tools, MCP configs, and agent skills on local machines — Initial-Copy332 · 2026-07-21
- A policy question: should output token poisoning by humans or AI be illegal? — slashML · 2026-07-21
- Open-source MCP proxy mcp-guard blocks prompt injection before tool calls run — TastePrestigious4419 · 2026-07-21
- Cancer-support-hub exposes 585+ cancer resources through an MCP connector — modelcontextprotocol · 2026-07-21
- Google DeepMind launches Gemini 3.5 Flash Cyber in a limited government-only pilot — ShakeelHashim · 2026-07-21
- Google launches Gemini 3.5 Flash Cyber, a cheaper AI model for vulnerability hunting — The Verge AI · 2026-07-21