Multi-teacher on-policy distillation (MOPD) explained: merging specialist LLMs into one student
cwolferesearch · x · 2026-10-07
Cameron Wolfe breaks down Multi-Teacher On-Policy Distillation (MOPD), a technique for consolidating several specialized teacher models into one student.
- OPD basics: the student generates rollouts, and the teacher scores each student token via a reverse-KL objective — dense supervision instead of one reward per trajectory as in outcome-based RL.
- Multi-teacher extension: rollouts are routed to the teacher specialized in that domain — SWE rollouts scored by a SWE teacher, chat prompts by a conversational teacher.
- Specialist teachers: teachers are trained independently with bespoke pipelines (SFT/RLVR/RLHF); Nemotron 3 Ultra trains 10 teacher models. Teachers whose behavior drifts too far from the student weaken supervision, fixed via MOPD warmup (SFT on teacher rollouts first).
- Why it matters: scaling RL across many domains suffers from rollout dilution, domain interference, and cost mismatches; distillation can be combined with policy-gradient RL rather than replacing it.
Related event: Multi-Teacher On-Policy Distillation Lets One Student Surpass Its Teachers(2 posts)→
More from Models
- Dev warns OpenRouter share, cache hit rate, latency stats are easily gamed for marketing — charles_irl · 2026-10-07
- OpenAI's unreleased model reportedly proves quasi-Riemann hypothesis with Lean proof — ChrisGPT · 2026-10-07
- Fed Claude Opus my blurry handheld Saturn shots, it fused them into one best image — adonis_singh · 2026-10-07
- Two Labs, One Race: Anthropic vs OpenAI Frontier Model Release Timeline, 2023–2026 — Medical-Sky7620 · 2026-10-07
- Claude was given robot skin — Opus was curious but anxious about hooking up — repligate · 2026-10-07
- Gemini 2.5 Pro retiring October 20, 2026, users say goodbye — hargup13 · 2026-10-07