Researchers argue catastrophic forgetting drives why fine-tuned bad behaviors persist in stronger models

QuintinPope5 · x · 2026-09-29

Quintin Pope argues that fine-tuning methods for removing bad model behaviors largely rely on catastrophic forgetting, which explains why such behaviors persist more strongly in larger models. He notes the setup — 'train to be bad on prefix x, good otherwise' — is nearly equivalent to simultaneously training bad behavior on x and good behavior elsewhere, making it unlikely the trained behaviors truly get removed.

Related event: Researcher questions whether catastrophic forgetting can erase LLM backdoor behaviors(4 posts)→

Original post →

More from Research

Research channel →