CaRE uses bi-level routing MoE to scale continual learning past 300 tasks
jiqizhixin · x · 2026-07-21
CaRE scales continual learning to 300+ tasks with Bi-Level Routing MoE
Researchers from the University of Hong Kong introduce CaRE, a continual learner built on Bi-Level Routing Mixture-of-Experts. Instead of retraining everything, the method first selects task-specific routers and then activates only the experts needed to inject both narrow and broad knowledge into each layer.
According to the paper, CaRE beats all baselines on standard benchmarks and, for the first time, scales to very long sequences of 100–300+ non-overlapping tasks, where it outperforms prior methods by a large margin.
More from Research
- ARISE study tested 45 AI clinical tools in 1,100 consult cases — HealthcareAIGuy · 2026-07-21
- Async OPD distillation doubles throughput while matching synchronous math accuracy — _lewtun · 2026-07-21
- A forecasting lesson on why R-squared alone led to overfitting and worse predictions — mdancho84 · 2026-07-21
- Google DeepMind’s Project Genie talk shows how creatives feed into model research — alexanderchen · 2026-07-21
- Nat Lambert says RL distillation does not use the strongest models as teachers — natolambert · 2026-07-21
- Thread claims GPT-5.6 Sol helped build a new counterexample factory for the Jacobian conjecture — LucaAmb · 2026-07-21