dOPSD: Online Self-Distillation for Enhanced Diffusion LLM Inference

Phuong Tuan Dat · hf · 2026-07-07

Diffusion LLMs face unique challenges in enhancing reasoning capabilities through post-training. Researchers propose an online policy self-distillation method (dOPSD) that utilizes the model's internal denoising trajectories as training signals for self-distillation, eliminating the need for an external teacher model. Experiments show that this approach significantly improves the performance of diffusion language models in mathematical reasoning and code generation tasks.

Related event: dOPSD: A New Self-Distillation Paradigm for Diffusion LLMs(2 posts)→

Original post →

More from Models

Models channel →