Composing Continual Learning Mechanisms Boosts 100-Task Memorization Retention 28x to 34.9%
DanielKhashabi · x · 2026-09-09
- New research on "long-horizon memorization": models learn 100 query-answer tasks via sequential SFT without retaining earlier examples or task IDs at inference.
- Naive sequential fine-tuning causes catastrophic forgetting (1.2% retention after 100 tasks); no single continual learning mechanism maintained strong retention at this horizon.
- The authors organize compositions along two dimensions: data/function/weight anchors (what prior information each update preserves) and low-rank allocation rules (where successive updates live).
- They use task-level successive halving to search the combinatorial space plus a factorial experiment to measure individual and interaction effects, validated on three distinct 100-task datasets.
- Best method combines all three anchors with merged LoRA: top-3 on all datasets, raising average final retention from 1.2% to 34.9% (28x). Data anchors and merged LoRA give the largest gains and interact super-additively across all three datasets.
More from Research
- AI researcher: benchmarks without released training data are 100% meaningless — mjdramstead · 2026-09-09
- antirez rebuts Terence Tao: AI proofs add to math knowledge, not subtract from it — antirez · 2026-09-09
- Flow explained in one Manim animation: trajectory, vector field, ODE — ariG23498 · 2026-09-09
- OVIE trains novel-view synthesis on 30M unpaired web images, runs 600x faster than rivals — ducha_aiki · 2026-09-09
- Bacterial collectives can track and report Xenopus embryos, Levin lab preprint shows — MacrinePhD · 2026-09-09
- Chess Study: When Players Spend Thinking Time Is Key to Cheating Detection — TZahavy · 2026-09-09