Researcher questions whether catastrophic forgetting can erase LLM backdoor behaviors
Quintin Pope argues that methods like BEEAR, which remove safety backdoors in LLMs, essentially rely on catastrophic forgetting, explaining why bad behaviors persist more in stronger models; he adds that LLM robustness to data ordering makes trained backdoor behaviors hard to erase.
2026-09-29 ~ 2026-09-29 · 4 related posts
- Quintin Pope: backdoor-style training setups are a poor stand-in for hypothesized inner optimizers — QuintinPope5 · 2026-09-29
- Quintin Pope: training's robustness to data ordering explains why backdoor behaviors survive — QuintinPope5 · 2026-09-29
- Researchers argue catastrophic forgetting drives why fine-tuned bad behaviors persist in stronger models — QuintinPope5 · 2026-09-29
- Safety backdoor removal via BEEAR: 95%+ to <1% attack success, but caveats for worst-case inner optimizers — QuintinPope5 · 2026-09-29