OPD² distills only the reasoning gains added by tuning, not the teacher’s old style
coallaoh · x · 2026-07-23
- The post highlights On-Policy Delta Distillation (OPD²), a new distillation method for reasoning models.
- Unlike standard on-policy distillation, which trains the student to imitate the teacher’s full output distribution, OPD² focuses on the delta signal: the difference between the teacher and its base model after reasoning tuning.
- The goal is to distill only the capability added by reasoning tuning, rather than also copying the teacher’s pre-existing style and preferences.
- The attached figure contrasts the two setups and frames OPD² as a way to preserve the learning trajectory that produces reasoning improvements.
More from Research
- Hugging Face releases The Stack v3, a 5T-token open code dataset — lvwerra · 2026-07-23
- Mila Quebec will host a talk on what AI benchmarks really measure for African languages — hugo_larochelle · 2026-07-23
- An agent stack diagram says production AI is 90% architecture, not prompts — theomitsa · 2026-07-23
- MIT’s free “SLAM for Dummies” guide turns robotics navigation into a hands-on tutorial — lukas_m_ziegler · 2026-07-23
- Sol 5.6 reportedly writes a full research paper from one prompt — conitzer · 2026-07-23
- NeurIPS position paper reviews are now out — hiddenmarkov · 2026-07-23