Local Support Learning: 7B LLMs learn new tasks at full capacity without forgetting or old data
CatAstro_Piyush · x · 2026-10-02
A new arXiv paper, Local Support Learning (LSL), proposes a general post-training framework against catastrophic forgetting: LLMs up to 7B parameters can learn new tasks at full capacity while retaining prior capabilities—without access to any prior data.
Key ideas:
- Casting forgetting as a geometric problem in the input space of each weight matrix, the authors show updates from standard gradient-based optimizers are suboptimal under the natural retention objective.
- LSL pairs a standard weight adapter (trained as usual) with a gating function that activates the adapter only on input activations from its own training distribution, making each update local to that distribution.
- The gate must route data from all learning phases while training only on current-phase data; the authors solve this with a Gaussian Mixture Model gate whose likelihood decays rapidly away from its training data, naturally staying closed on prior-phase inputs.
Experiments show the post-training approach retains both pretrained and finetuned capabilities across multiple training phases, with low memory/compute overhead, robustness to hyperparameter choice, and scaling potential. Paper and code are public.
More from Research
- SFT is not dead: sampling-rewritten data rivals RL posttraining, with better generalization — mayfer · 2026-10-02
- Where-OPD: synthetic-scene spatial self-distillation boosts MLLM perception by 3.23 points — valeocorg · 2026-10-02
- Stability AI's SemanTok: 201M AR video model matches a 3.4x larger rival with semantic tokens — stabilityai · 2026-10-02
- An 'Amazon's Choice' label flips LLM picks: three NeurIPS 2026 bias papers — xuandongzhao · 2026-10-02
- Prompts revealed that make AI text score 100% human on detectors — paulnovosad · 2026-10-02
- Pangram intentionally avoids flagging AI-translated texts as AI-generated — paulnovosad · 2026-10-02