Why do research labs prefer Off-Policy Distillation for model improvement?
miifanboy · reddit · 2026-08-28
The author questions why empero-ai used Off-Policy Distillation to distill Qwen3.8 2.4T A95B into older Qwen3.5 releases, arguing that On-Policy Distillation would yield better results. The author believes On-Policy methods allow the actual KL divergence to match the teacher model better, truly distilling knowledge rather than just cloning behavior and hoping for the best.
More from Models
- Planned Rerun of 3D Game Benchmark for Grok 4.6 and GLM 5.3 Flash — kevinkern · 2026-08-28
- Local LLMs Still Lagging Behind Frontier Models — sdmat123 · 2026-08-28
- Own a frontier AI model running locally in just 5 hours — MaziyarPanahi · 2026-08-28
- Qwen3.8-Flash-Next Released: A Free AI Rivaling Billion-Dollar Giants — Two Minute Papers · 2026-08-28
- Users suspect DeepSeek V4 quality drop due to routing changes — l33thax0r_ · 2026-08-28
- DeepSeek V4 Flash Pricing Matches GPT OSS 20B, Sparking Cost Discussion — ChrizBogota · 2026-08-28