Weak-to-Strong Distillation Paper Shows Qwen3-8B Surpassing Its 4B Teacher

burny_tech · x · 2026-07-31

A new paper, "Weak-to-Strong On-Policy Distillation," introduces a novel approach to model distillation.

Strong models typically require even stronger teachers for OPD, which are hard to find for frontier models. This paper demonstrates that a strong model can be improved using weaker models by distilling the logit direction between a weak positive and a weaker negative model, rather than copying either one. In their results, Qwen3-8B showed significant improvements in math and code, occasionally surpassing the 4B RL teacher model it learned from.

Original post →

More from Research

Research channel →