New Self-Distillation Paradigm for Diffusion LLMs
机器之心 · wechat · 2026-07-09
This research, conducted by institutions including the Max Planck Institute and Tsinghua University, proposes d-OPSD for on-policy self-distillation in diffusion large language models. The paper explains that d-OPSD eliminates the need for reference solutions or additional teacher models. Instead, the student model samples online first, and its trajectories are fed back to the teacher as privileged information. Experiments show that across multiple mathematical reasoning benchmarks, this method achieves or surpasses RL performance using fewer training steps.
Related event: dOPSD: A New Self-Distillation Paradigm for Diffusion LLMs(2 posts)→
More from Research
- OpenForecaster uses daily news to improve language-model forecasting — Cohere_Labs · 2026-07-21
- SenseTime unveils U1 Pro and open-sources a 50M-sample vision dataset at WAIC 2026 — 机器之心 · 2026-07-21
- Baseten study finds new facts in LLM weights are fragile unless trained from many restatements — alex_verem · 2026-07-21
- Kimi K3 and Fable 5 now look much closer than the old open-vs-closed gap — FinanceYF5 · 2026-07-21
- uv-scripts/ocr returns to the top of Hugging Face datasets with a JSON model picker — vanstriendaniel · 2026-07-21
- DeepSearch-World trains web agents with 420K verifiable QA tasks — HKUST · 2026-07-21