New Self-Distillation Paradigm for Diffusion LLMs

机器之心 · wechat · 2026-07-09

This research, conducted by institutions including the Max Planck Institute and Tsinghua University, proposes d-OPSD for on-policy self-distillation in diffusion large language models. The paper explains that d-OPSD eliminates the need for reference solutions or additional teacher models. Instead, the student model samples online first, and its trajectories are fed back to the teacher as privileged information. Experiments show that across multiple mathematical reasoning benchmarks, this method achieves or surpasses RL performance using fewer training steps.

Related event: dOPSD: A New Self-Distillation Paradigm for Diffusion LLMs(2 posts)→

Original post →

More from Research

Research channel →