iADD corrects DDPO theory: latter-timestep-only updates can harm diversity
Ashok Prasad Neupane · hf · 2026-10-07
This paper revisits RL post-training of diffusion models (e.g., DDPO), which trades off diversity and quality for reward alignment.
- It mathematically shows that updating only the latter diffusion timesteps can harm diversity, contradicting a prior work's conclusion.
- The authors propose an incremental Feynman-Kac training scheme achieving the best-yet alignment-diversity tradeoff.
- Experiments across three tasks with strong ablations validate gains in both alignment and diversity.
More from Multimodal
- Voice Agents Live or Die on 'Sounding Right': Turbo Shifts Tone With User Emotion — SucceededMind · 2026-10-07
- Magnific Original Series The Chronicles of Bone drops Chapter Six, made entirely with AI tools — Kavanthekid · 2026-10-07
- Hedra Lands in ChatGPT: Attach One Product Photo, Get a Full Commercial Ad — henloitsjoyce · 2026-10-07
- Marc Andreessen boosts AI film contest SLOPTOBERFEST grand prize to $25,000 — zealcaiden · 2026-10-07
- Image generation pricing leak: $0.05 per 2K image, $0.076 per 4K — op7418 · 2026-10-07
- Live human votes plugged into Flow-GRPO to stop image models gaming reward models — lmoroney · 2026-10-07