Researchers argue catastrophic forgetting drives why fine-tuned bad behaviors persist in stronger models
QuintinPope5 · x · 2026-09-29
Quintin Pope argues that fine-tuning methods for removing bad model behaviors largely rely on catastrophic forgetting, which explains why such behaviors persist more strongly in larger models. He notes the setup — 'train to be bad on prefix x, good otherwise' — is nearly equivalent to simultaneously training bad behavior on x and good behavior elsewhere, making it unlikely the trained behaviors truly get removed.
More from Research
- TT-VidT: temporal-decoupled video pretraining accepted at NeurIPS — CMHungSteven · 2026-09-29
- EgoDemo, an egocentric human demonstration dataset for embodied AI, trends on HF — LightwheelAI · 2026-09-29
- Peking University: post-training leaves behavioral shadows, one word per prompt transfers coding skill — PekingUniversity · 2026-09-29
- RenderRank reranks documents as images: 35% fewer tokens, beats sub-4B text rerankers — nlpai-lab · 2026-09-29
- DepthBench: residual connection design decides whether depth is a real scaling axis — Intelligent-Systems · 2026-09-29
- DN-MOPD: domain-normalized feedback fixes multi-teacher on-policy distillation imbalance — NanyangTechnologicalUniversity · 2026-09-29