MOPD Distillation in MiMo-V2-Flash

青稞AI · wechat · 2026-07-14

This note focuses on MOPD (Multi-Teacher On-Policy Distillation) from the MiMo-V2-Flash Technical Report, explaining how it reframes multi-expert capability integration into a single student model via on-policy reinforcement learning.

Background: Why MOPD?

Core Mechanism

Role in the Training Pipeline

The report divides post-training into three stages:

The author notes that MoE SFT stability depends on hyperparameters like num-zeros and expert bias update rates. Domain-specific RL, particularly for code agent training, was conducted at a massive scale using containerized environments and large-scale concurrent clusters.

Mathematical Formulation & Engineering Details

Experimental Conclusions

Original post →

More from Infra

Infra channel →