Study Finds Weak Same-Origin Teachers Outperform Strong Cross-Origin Ones in LLM Distillation
rohanpaul_ai · x · 2026-08-30
This study on On-Policy Distillation (OPD) reveals that student models copy their teacher's reasoning behavior rather than specific answers.
Key Findings:
- Lineage Matters More Than Benchmarks: A weaker teacher from the same base model family often outperforms a stronger teacher from a different family.
- Data Is Secondary: Training difficulty barely impacts generalization; even problems the teacher fails to solve are useful.
- Double-Edged Sword: Same-origin distillation enables transfer across languages and domains, but mixing multi-source teachers causes a "seesaw" effect where capabilities fluctuate based on the mixture.
More from Research
- Strong Models Design Harnesses for Weak Ones: Accuracy Nearly Doubles Without Training — 机器之心 · 2026-08-30
- Learn Positional Encodings derivation from first principles — zainhas · 2026-08-30
- COLM Paper Traces Capability Provenance in LLMs via Gradient Attribution — ziv_ravid · 2026-08-30
- Toby Ord paper argues recursive self-improvement has physical limits — Exponential View (Azeem Azhar) · 2026-08-30
- AI Formalization Tools Fable and Sol Spot First Repairable Error in Published Literature — Sauers_ · 2026-08-30
- Mark Schmidt Posts ICML Tutorial Video: Is Numerical Optimization Theory Irrelevant to ML Practice in 2026? — MarkSchmidtUBC · 2026-08-30