DOPD: Novel On-Policy Distillation to Tackle Privilege Hallucination
jiqizhixin · x · 2026-07-16
Researchers from NUS, CUHK, Peking University, and JD.com proposed a new LLM distillation method called **DOPD** (Dual On-policy Distillation), aiming to solve the "privilege hallucination" problem in traditional distillation. Instead of blindly copying the teacher model's hidden advantages, this method dynamically decides which tokens to extract for supervision from either the teacher or the student's own policy. This prevents the model from confusing knowledge it can actually learn with abilities it can only "fake." Experiments show that DOPD outperforms standard on-policy distillation on both LLM and VLM benchmarks, achieving additional improvements in stability, robustness, continual learning, and out-of-distribution tasks.
More from Research
- Draft paper uses Markov-chain eigenfunctions to build partitions and speed up sampling — michaelchchoi · 2026-07-21
- Autoresearch proposes packaging ML runs as studies with questions, analysis, and code diffs — morgymcg · 2026-07-21
- GitHub repo adds lightweight ternary QAT for Prism-ML Bonsai models — terminoid_ · 2026-07-21
- Qdrant co-hosts a Munich meetup on search, retrieval, and agentic RAG on July 23 — qdrant_engine · 2026-07-21
- GigaChat Audio targets long-form audio grounding with timestamps across 120-minute inputs — ai-sage · 2026-07-21
- Paper models Transformer components as stochastic geometry and tests five architectures — Zhihua Liang · 2026-07-21