MOPD-Router: token-level teacher routing boosts multi-teacher distillation by up to 12.3%
GAIR · hf · 2026-09-28
GAIR introduces MOPD-Router, a framework that routes supervision across the full teacher pool at each token in multi-teacher on-policy distillation, without domain labels or a separate routing model.
- Problem: existing prompt-level hard routing wastes complementary signals and requires domain labels.
- ExpertAlign: scores each teacher by whether its token-level correction reflects its post-training specialization, beating Entropy- and Novelty-based metrics.
- Results: best in all four settings; +5.88 points (+12.3%) over Mean aggregation on unlabeled mixtures, +3.95 points (+7.8%) over standard MOPD on labeled data without using labels.
Code is open-sourced on GitHub.
More from Research
- Simon Prince Publishes Neural ODEs Tutorial: Residual Nets as Infinite-Layer ODEs — SimonPrinceAI · 2026-09-28
- Xiaomi open-sources 7,000+ RL environments with an interactive HF Space to run rollouts — SergioPaniego · 2026-09-28
- 77.8% of Agent Runs Carry Stale File Views — Lessons From Building a Context-Refresh Layer — jabulari · 2026-09-28
- GPT Researcher drops embeddings: LLM-based retrieval lifts relevant context 59% at same cost — hwchase17 · 2026-09-28
- Meta claims it can fix AI slop with RL-XAR: RL on expert-aligned rubrics for writing — jaseweston · 2026-09-28
- Policy-DRIFT lands NeurIPS acceptance with 49% drag reduction, beating DRL by 16% — ricardovinuesa · 2026-09-28