Study Finds On-Policy Distillation Transfers Reasoning, Not Answers
teortaxesTex · x · 2026-08-30
This paper investigates the generalization behavior of On-Policy Distillation (OPD). Key findings include:
- Training difficulty has little effect: Problems the teacher never solves still provide useful supervision, suggesting OPD transfers the teacher's reasoning behavior rather than specific answers.
- Model origin matters: Same-origin teacher-student pairs generalize well across languages, reasoning horizons, and domains, while cross-origin teachers often transfer less effectively.
- Multi-teacher trade-offs: Routing prompts to domain experts fails to isolate influence; changing the teacher mixture causes a seesaw effect in student capabilities rather than simple skill combination.
More from Research
- Strong Models Design Harnesses for Weak Ones: Accuracy Nearly Doubles Without Training — 机器之心 · 2026-08-30
- Learn Positional Encodings derivation from first principles — zainhas · 2026-08-30
- COLM Paper Traces Capability Provenance in LLMs via Gradient Attribution — ziv_ravid · 2026-08-30
- Toby Ord paper argues recursive self-improvement has physical limits — Exponential View (Azeem Azhar) · 2026-08-30
- AI Formalization Tools Fable and Sol Spot First Repairable Error in Published Literature — Sauers_ · 2026-08-30
- Mark Schmidt Posts ICML Tutorial Video: Is Numerical Optimization Theory Irrelevant to ML Practice in 2026? — MarkSchmidtUBC · 2026-08-30