DN-MOPD: domain-normalized feedback fixes multi-teacher on-policy distillation imbalance

NanyangTechnologicalUniversity · hf · 2026-09-29

NTU researchers study multi-teacher on-policy distillation (MOPD), where RL-trained single-skill specialists jointly teach one student. They find feedback imbalance—instruction-following feedback is several times more spread out than math feedback and dominates updates—so the student gains little of the math specialist's edge.

Takeaway: merging specialists requires deciding not only who teaches but how strongly.

Original post →

More from Research

Research channel →