Nemotron 3 Ultra report reveals MOPD distillation teachers must share compatible training pipelines

cwolferesearch · x · 2026-10-02

Key finding

Drawing from NVIDIA's Nemotron tech reports, the author highlights a subtle but important lesson in Multi-Teacher On-Policy Distillation (MOPD): not all teacher models can be combined effectively.

How MOPD works

Two main setups

The lesson

The Nemotron 3 Ultra report states that teachers trained with substantially different pipelines cannot be effectively merged via straightforward MOPD. Student-generated trajectories become out-of-distribution for such teachers, making token-level supervision far less useful. This mainly hurts the consolidation setup; recovery is naturally safe since teachers come from the same lineage.

Fix

MOPD warmup: a short SFT stage on diverse teacher-sampled trajectories before distillation moves the student closer to teacher behavior. Bottom line: MOPD isn't just about distilling experts—student/teacher compatibility is essential.

Original post →

More from Research

Research channel →