Policy Gradient in Diffusion Models Sparks Debate

Developers discussed the high cost of end-to-end backpropagation in diffusion models, noting that policy gradients are rarely used. ByteDance's VLM reward paper might be an exception, utilizing a branch rollback abstraction typically necessary only for long-range language models.

2026-08-02 ~ 2026-08-02 · 3 related posts