DMSampler Accelerates Diffusion RL Training, Cutting GPU Hours by 10x

jiqizhixin · x · 2026-08-12

Researchers from USTC and collaborators introduced DMSampler, a new method to tackle the high computing power consumption in training image and video generation models.

The approach replaces the slow traditional 50-step sampling with a fast 4-to-8-step distilled model acting as a quick proxy for the AI policy. It continuously updates alongside the policy to ensure samples remain accurate and aligned throughout training.

Experiments show that DMSampler outperforms previous diffusion RL methods on OCR, GenEval, and VBench benchmarks while reducing GPU hours by an order of magnitude.

Original post →

More from Multimodal

Multimodal channel →