DiffusionOPSD: On-Policy Self-Distillation in Diffusion Models

ByteDance-Seed · hf · 2026-08-26

ByteDance proposed DiffusionOPSD, which uses on-policy self-distillation to turn image-level rewards into explicit intermediate targets for diffusion models. This approach improves alignment efficiency and enables separate analysis of target construction and policy fitting.

Original post →

More from Research

Research channel →