iADD corrects DDPO theory: latter-timestep-only updates can harm diversity

Ashok Prasad Neupane · hf · 2026-10-07

This paper revisits RL post-training of diffusion models (e.g., DDPO), which trades off diversity and quality for reward alignment.

Original post →

More from Multimodal

Multimodal channel →