DN-MOPD: domain-normalized feedback fixes multi-teacher on-policy distillation imbalance
NanyangTechnologicalUniversity · hf · 2026-09-29
NTU researchers study multi-teacher on-policy distillation (MOPD), where RL-trained single-skill specialists jointly teach one student. They find feedback imbalance—instruction-following feedback is several times more spread out than math feedback and dominates updates—so the student gains little of the math specialist's edge.
- DN-MOPD keeps routing but rescales each domain's feedback by its measured spread
- Across Qwen3.5 at three sizes, six benchmarks, three seeds and two length caps, DN-MOPD beats MOPD and recovers most of the lost math gain
- Controls show gains come mainly from turning down instruction-following feedback; fixed weights near measured values perform comparably
Takeaway: merging specialists requires deciding not only who teaches but how strongly.
More from Research
- ColNanoVDR distills multi-vector visual document retrieval without documents, keeping 95% NDCG@5 at 149M params — nanovdr · 2026-09-29
- NUS rethinks DiT residual connectivity: 1.73x fewer training iterations, 1.39 FID — NationalUniversityofSingapore · 2026-09-29
- Imprint Reader Decodes Weight Updates into Natural Language, Enables Targeted Edits — Guanxu Chen · 2026-09-29
- SJTU's GeoVerse Synthesizes World-Consistent Novel Views in Geometric Latent Space — SJTU · 2026-09-29
- Tencent Hunyuan Maps Scaling Laws for Encoder-Free Multimodal Pretraining — Tencent-Hunyuan · 2026-09-29
- When Do Model Internals Help? Benchmarking Representation Engineering for LLM Safety — Tianyi Guan · 2026-09-29