MIT's Local Support Learning Fixes Catastrophic Forgetting in LLMs Without Old Data
MIT · hf · 2026-10-05
MIT researchers propose Local Support Learning (LSL), a general post-training framework that prevents catastrophic forgetting without access to prior data.
- Frames forgetting as a geometric problem in each weight matrix's input space, showing gradient-based updates are suboptimal for retention
- LSL pairs a standard weight adapter with a gating function that activates the adapter only on inputs from its own training distribution, localizing each update
- The gate is a Gaussian Mixture Model whose likelihood decays rapidly away from training data, naturally staying closed on earlier-phase inputs
- Resolves forgetting in LLMs up to 7B parameters across multiple training phases, retaining both pretrained and finetuned capabilities, while being memory/compute-efficient and robust to hyperparameter choices
More from Research
- Human-AI collaboration pushes 11-square packing lower bound to 3.875, closing 95.89% of the gap — ctjlewis · 2026-10-05
- Follow-up on the same 11-square packing breakthrough: 300-line verifiable proof via 1,121 weighted dots — ctjlewis · 2026-10-05
- Pivot-SD: 200 questions beat diffusion RL by distilling only high-impact denoising pivots — kaist-ai · 2026-10-05
- SBERT2S1 turns biomedical retrieval encoders into calibrated one-pass typed decision models — Pritam Deka · 2026-10-05
- 315M-param self-play RL bot Pluto crushes GPT-6 Astra 0-1000 in StarCraft: Brood War — i_dg23 · 2026-10-05
- Same system scores 51.4 vs 75.8 depending on judge — are memory benchmarks trustworthy? — Efficient_Joke3384 · 2026-10-05