d-OPSD: Self-Distillation from Future Answers

jiqizhixin · x · 2026-07-17

This paper introduces d-OPSD, an on-policy self-distillation framework for diffusion LLMs. The core idea: instead of only learning from past tokens, the model learns from its own generated "future answers," an approach the authors term suffix conditioning.

Method Highlights

Experimental Conclusions

The paper is titled Learning from the Self-future: On-policy Self-distillation for dLLMs.

Original post →

More from Research

Research channel →