NAVER Proposes Delta Distillation Method
naver-ai · hf · 2026-07-20
NAVER AI Lab proposed **On-Policy Delta Distillation (OPD²)**, shifting the distillation objective from "imitating the teacher's output distribution" to supervising the **delta signal** between the teacher model and its base model—essentially the changes induced by instruction tuning. The authors argue that this signal transfers reasoning capabilities more directly. Experiments covering math, science, and code reasoning benchmarks demonstrate that OPD² is more stable and effective than traditional on-policy distillation, yielding stronger reasoning model performance within a shorter post-training phase. Code will be open-sourced.
More from Research
- Draft paper uses Markov-chain eigenfunctions to build partitions and speed up sampling — michaelchchoi · 2026-07-21
- Autoresearch proposes packaging ML runs as studies with questions, analysis, and code diffs — morgymcg · 2026-07-21
- GitHub repo adds lightweight ternary QAT for Prism-ML Bonsai models — terminoid_ · 2026-07-21
- Qdrant co-hosts a Munich meetup on search, retrieval, and agentic RAG on July 23 — qdrant_engine · 2026-07-21
- GigaChat Audio targets long-form audio grounding with timestamps across 120-minute inputs — ai-sage · 2026-07-21
- Paper models Transformer components as stochastic geometry and tests five architectures — Zhihua Liang · 2026-07-21