Multi-Teacher On-Policy Distillation Lets One Student Surpass Its Teachers
Multi-Teacher On-Policy Distillation (MOPD) merges multiple specialist teachers into one student model; the new Latent-MOPD extends it to the representation level using hidden states, with a single student outperforming each teacher.
2026-10-06 ~ 2026-10-07 · 2 related posts
- Latent-MOPD: Representation-Level Multi-Teacher Distillation Lets One Student Beat Every Teacher — zillow · 2026-10-06
- Multi-teacher on-policy distillation (MOPD) explained: merging specialist LLMs into one student — cwolferesearch · 2026-10-07