MOPD drops OPSD's self-teaching flaw, but naive multi-teacher averaging may distill style over capability
novasarc01 · x · 2026-09-28
The author argues MOPD (multi-teacher on-policy distillation) removes a troubling feature of OPSD: the student no longer has to 'self-teach' from privileged context, and can instead learn from genuinely stronger, domain-specific teachers.
But this creates a bias-selection problem: naively averaging teacher losses may distill teacher style and inter-teacher disagreement rather than real capability.
His proposed direction is 'selective MOPD': dynamically weight or gate teacher gradients based on whether they're likely to improve downstream reward — provided that can be estimated cheaply online without massive counterfactual rollout costs. He sees this as a much stronger path than vanilla multi-teacher distillation.
More from Research
- Inference Engineering Archive launches: end-to-end resource from first principles to production — soham_btw · 2026-09-28
- OpenAI's thousands of agents solved Navier-Stokes; mathematicians worry brute force erases the art — nordicinst · 2026-09-28
- University of Tokyo grows self-healing living human skin on robotic finger — TinfoilTricorn · 2026-09-28
- Thales refines satellite DSMs with diffusion models, cutting urban RMSE from 6.00 to 3.45 m — thalesgroup · 2026-09-28
- Stanford's HomeBody gives frontier VLMs spatial memory and composable humanoid skills — CyberRobooo · 2026-09-28
- UK Startup Turns Gaming Inputs Into Training Data for World Models, With LeCun Advising — nordicinst · 2026-09-28