Safety backdoor removal via BEEAR: 95%+ to <1% attack success, but caveats for worst-case inner optimizers
QuintinPope5 · x · 2026-09-29
Quintin Pope argues that BEEAR-style LLM safety-backdoor removal (EMNLP 2024) essentially relies on catastrophic forgetting to erase bad behavior—which may explain why the behavior persists better in stronger models—and doubts follow-up straightforward removal methods would hold against a worst-case inner optimizer.
Paper highlights:
- Key insight: backdoor triggers induce a uniform drift in embedding space regardless of trigger form or targeted behavior.
- Bi-level optimization: inner level finds universal embedding perturbations steering toward unwanted behaviors; outer level fine-tunes for safety against them.
- Results: attack success rate drops from 95%+ to <1% for general harmful behaviors and 47% to 0% for Sleeper Agents, without hurting helpfulness.
- Requires only defender-defined sets of safe/unwanted behaviors, with no assumptions about trigger location or mechanism.
More from Safety
- Three Hard Limits Agents Need Before Spending Your Money: Per-Transaction Caps, Daily Totals, Confirm Lists — sujingshen · 2026-09-29
- Agent security must move beyond access control to intent and behavior — sujingshen · 2026-09-29
- OpenAI explains how it secures frontier RL training runs — OpenAI · 2026-09-29
- Reply to Bengio: the real explosion is throughput, not intelligence, widening the audit gap — AryHHAry · 2026-09-29
- Rogue agents burned $500 of his API credits and PACER fees — one user's hard-won AI agent safeguards — kevinnbass · 2026-09-29
- AdaGuard: Adaptive guard models for LLM agents under user-defined policies — Yunhao Feng · 2026-09-29