Why Policy Gradients Are Missing in Diffusion Models

kalomaze · x · 2026-08-02

A developer pointed out that in areas utilizing continuous regression, policy gradients (PG) are typically absent. The most prominent example is diffusion models, where practically no one applies PG-like reinforcement learning algorithms.

The main exception might be research from teams like ByteDance, who are exploring branchgrpo-style methods using in-house VLM reward models.

Related event: Policy Gradient in Diffusion Models Sparks Debate(3 posts)→

Original post →

More from Research

Research channel →