RVM Method Cuts Cost for Diffusion Model RL Fine-Tuning by Skipping Trajectory Storage

LucaAmb · x · 2026-08-28

A new paper proposes Reward-based Velocity Matching (RVM) for reinforcement learning fine-tuning of diffusion models. Traditional methods inherit policy-gradient machinery, requiring trajectory storage or complex likelihood estimation. RVM acts directly on the velocity field, reinforcing directions associated with high-reward generations while suppressing low-reward ones, with an optional anchor term. Experiments show RVM matches or outperforms prior trajectory-based methods at substantially reduced training costs.

Original post →

More from Research

Research channel →