KL Divergence Isn't Needed for On-Policy Distillation, Only Token Update Direction Matters
Wenze Lin · hf · 2026-09-30
- Key finding: KL divergence is unnecessary for on-policy distillation (OPD) — preserving the update direction toward the teacher suffices. Critically, only a small subset of tokens with strong teacher-student disagreement matters; other tokens can even be pulled away from the teacher.
- Simplified loss: Assigning +1 to tokens where teacher probability exceeds the student and -1 otherwise reproduces nearly the same training behavior as reverse-KL OPD.
- Application — C-MOPD: Unlike MOPD's per-sample single-teacher routing (which causes capability conflicts), C-MOPD supervises every sample with all teachers, consistently beating MOPD on math and code benchmarks.
- Code: github.com/LeapLabTHU/KL-Free-OPD.
More from Research
- Gary Marcus doubles down on 2020 thesis: LLMs alone aren't enough for robust AI — GaryMarcus · 2026-09-30
- DepthBench paper finds Pre-LN variants hit a depth wall, comparing 10 residual designs — teortaxesTex · 2026-09-30
- Xcelsa's Apex AI fixes chip timing violation in 3 hours vs 1 month — IanAndrewsDC · 2026-09-30
- Fatima Fellowship opens first Fall cohort after 700+ spring applicants — deliprao · 2026-09-30
- flyverse: open framework runs whole fly connectomes in strange simulated worlds — repligate · 2026-09-30
- Daniel Litt: once models write well, human writing will mainly help you think — littmath · 2026-09-30