Same-family on-policy distillation scaling laws: smaller teachers can out-transfer bigger ones to strong students
zju · hf · 2026-10-01
This paper studies scaling properties of on-policy distillation (OPD) across weak-to-strong, same-base and strong-to-weak setups. Early training shows a consistent useful-transfer regime: held-out gold score rises roughly linearly in the square root of token-level reverse KL from the student's initialization. In every weak-to-strong pair, the student's peak score exceeds its teacher's own, so compact RL experts can transfer reasoning to much larger students. Fitted power laws show peak score gains flatten once teacher size approaches student size, and at matched gold score smaller teachers transfer better — a teacher's score alone doesn't define its supervision value.
More from Research
- Fudan and CUHK MMLab publish first survey of Agentic Visual Generation, classifying systems L0-L4 — jiqizhixin · 2026-10-01
- AI Loop Finds Rare Variant Human Labs Missed Due to Distance-Based Filtering — danielmckinn0n · 2026-10-01
- MICA: an open-source LM with no neural network, just learned cellular automata rules — Silver_Employ2617 · 2026-10-01
- UniReps Workshop Papers Due Oct 4: Model Merging, Representational Alignment and More — ClementineDomi6 · 2026-10-01
- Forked memory pages: keeping worlds separate so old knowledge never gets overwritten — RexDouglass · 2026-10-01
- DeepMind and Isomorphic Labs unveil bioresilience plan backed by 15+ partnerships — davidstutz92 · 2026-10-01