DiffusionOPSD Cuts Diffusion Model Training Compute by 63%

burny_tech · x · 2026-08-28

The paper 'On-Policy Self-Distillation in Diffusion Models' introduces DiffusionOPSD, addressing the inefficiency of Diffusion RL which provides rewards only at the final image. This method converts endpoint rewards into explicit reward-improving targets at intermediate states and distills them back into the model via on-policy self-distillation. Experiments show that across two diffusion backbones and ten evaluators, DiffusionOPSD achieved the best held-out score in 19 of 20 settings while cutting training GPU-hours by up to 63%. The core idea frames diffusion alignment as continuous self-distillation with direct intermediate supervision.

Related event: DiffusionOPSD Cuts Diffusion Model Training Compute by 63%(2 posts)→

Original post →

More from Infra

Infra channel →