ByteDance Seed's DiffusionOPSD: full-trajectory rewards for diffusion RL, 60% less GPU time

andrew_n_carr · x · 2026-08-29

ByteDance's Seed team released DiffusionOPSD (On-Policy Self-Distillation in Diffusion Models), tackling a core pain point of RL for diffusion models: endpoint rewards only arrive after the full image is decoded, leaving intermediate denoising steps unsupervised.

Method:

Results: Extensive ablations on SD3.5-M and step-distilled Z-Image-Turbo across ten evaluators; best final held-out scores in reward-matched settings with roughly 60% fewer training GPU-hours than DiffusionNFT. Paper and code are public.

Original post →

More from Research

Research channel →