MOPD drops OPSD's self-teaching flaw, but naive multi-teacher averaging may distill style over capability

novasarc01 · x · 2026-09-28

The author argues MOPD (multi-teacher on-policy distillation) removes a troubling feature of OPSD: the student no longer has to 'self-teach' from privileged context, and can instead learn from genuinely stronger, domain-specific teachers.

But this creates a bias-selection problem: naively averaging teacher losses may distill teacher style and inter-teacher disagreement rather than real capability.

His proposed direction is 'selective MOPD': dynamically weight or gate teacher gradients based on whether they're likely to improve downstream reward — provided that can be estimated cheaply online without massive counterfactual rollout costs. He sees this as a much stronger path than vanilla multi-teacher distillation.

Original post →

More from Research

Research channel →