Why do research labs prefer Off-Policy Distillation for model improvement?

miifanboy · reddit · 2026-08-28

The author questions why empero-ai used Off-Policy Distillation to distill Qwen3.8 2.4T A95B into older Qwen3.5 releases, arguing that On-Policy Distillation would yield better results. The author believes On-Policy methods allow the actual KL divergence to match the teacher model better, truly distilling knowledge rather than just cloning behavior and hoping for the best.

Original post →

More from Models

Models channel →