CaRE uses bi-level routing MoE to scale continual learning past 300 tasks

jiqizhixin · x · 2026-07-21

CaRE scales continual learning to 300+ tasks with Bi-Level Routing MoE

Researchers from the University of Hong Kong introduce CaRE, a continual learner built on Bi-Level Routing Mixture-of-Experts. Instead of retraining everything, the method first selects task-specific routers and then activates only the experts needed to inject both narrow and broad knowledge into each layer.

According to the paper, CaRE beats all baselines on standard benchmarks and, for the first time, scales to very long sequences of 100–300+ non-overlapping tasks, where it outperforms prior methods by a large margin.

Original post →

More from Research

Research channel →