Same-family on-policy distillation scaling laws: smaller teachers can out-transfer bigger ones to strong students

zju · hf · 2026-10-01

This paper studies scaling properties of on-policy distillation (OPD) across weak-to-strong, same-base and strong-to-weak setups. Early training shows a consistent useful-transfer regime: held-out gold score rises roughly linearly in the square root of token-level reverse KL from the student's initialization. In every weak-to-strong pair, the student's peak score exceeds its teacher's own, so compact RL experts can transfer reasoning to much larger students. Fitted power laws show peak score gains flatten once teacher size approaches student size, and at matched gold score smaller teachers transfer better — a teacher's score alone doesn't define its supervision value.

Original post →

More from Research

Research channel →