Safety backdoor removal via BEEAR: 95%+ to <1% attack success, but caveats for worst-case inner optimizers

QuintinPope5 · x · 2026-09-29

Quintin Pope argues that BEEAR-style LLM safety-backdoor removal (EMNLP 2024) essentially relies on catastrophic forgetting to erase bad behavior—which may explain why the behavior persists better in stronger models—and doubts follow-up straightforward removal methods would hold against a worst-case inner optimizer.

Paper highlights:

Related event: Researcher questions whether catastrophic forgetting can erase LLM backdoor behaviors(4 posts)→

Original post →

More from Safety

Safety channel →