Diffusion models can backprop end-to-end but costly; policy gradients only necessary for long-horizon language models

kalomaze · x · 2026-08-02

Author kalomaze notes that diffusion models can in principle backprop through the entire chain differentiably, but it's expensive. This cost forces the branching rollouts abstraction from ByteDance papers, which is only necessary for long-horizon language models. He also observes that almost no one uses policy gradients on diffusion models, except possibly ByteDance's branchgrpo-style works with in-house VLM reward models.

Related event: Policy Gradient in Diffusion Models Sparks Debate(3 posts)→

Original post →

More from Research

Research channel →