PaLoRA derives rank-aware pacing law for LoRA continual learning, +4% on 50-task benchmarks
Yuxuan Li · hf · 2026-10-06
A new paper, PaLoRA, tackles catastrophic forgetting in LoRA-based continual learning with a theoretical pacing rule.
- Finding: even under directional constraints like nullspace projection, finite-precision updates leak into prior-knowledge subspaces; small learning rates attenuate but cannot stop forgetting from intensifying as the effective rank of past knowledge grows.
- Theory: under an anisotropic leakage model, the optimal magnitude restriction scales as s = R/c, where R is the effective rank of past updates.
- Method: adaptive SVD truncation compresses historical knowledge, gradients are projected onto prior tasks' nullspace, and rank-aware adaptive pacing is applied.
- Results: consistent gains over prior methods, with up to 4% accuracy improvement on 50-task ImageNet-A and ImageNet-R in long-horizon settings.
More from Research
- Cornell Tech prof presents two pretraining-behavior papers at COLM, recruiting PhDs and postdocs — pratyushmaini · 2026-10-06
- Non-Expert Runs AI-Driven Interpretability Study Claiming Linear Representation of Directive Force in LLMs — LoudYogurtcloset7856 · 2026-10-06
- AC2 lets LLM RL train faster than GRPO by trusting critics on partial rollouts — stanfordnlp · 2026-10-06
- Stanford CS224N Winter 2026 posts free slides and assignments online — stanfordnlp · 2026-10-06
- Only right answers still teach chemistry: popular chemistry benchmark has shortcuts — tak3sh8 · 2026-10-06
- UT Austin math chair: OpenAI appears set to release ~400 AI-generated proofs at once — 141_1337 · 2026-10-06