Researcher questions whether catastrophic forgetting can erase LLM backdoor behaviors

Quintin Pope argues that methods like BEEAR, which remove safety backdoors in LLMs, essentially rely on catastrophic forgetting, explaining why bad behaviors persist more in stronger models; he adds that LLM robustness to data ordering makes trained backdoor behaviors hard to erase.

2026-09-29 ~ 2026-09-29 · 4 related posts