CompVis improves Distributional Diffusion Models: 4.48 FID at 4 steps on ImageNet

CompVis · hf · 2026-09-30

CompVis releases Improved Distributional Diffusion Models (iDDM). DDMs replace mean-prediction denoisers with distributional denoisers trained via a scoring rule, but faced two scaling obstacles: multi-particle training overhead and globally fixed scoring-rule hyperparameters forcing one trade-off across sampling budgets.

The improvements defer particle expansion to late transformer layers and introduce time-dependent scoring rule schedules informed by the dynamical regimes of Biroli 2024. With a DiT-based latent setup, a DiT-XL/2 trained from scratch in one stage (no teacher, self-distillation or JVPs) reaches 4.48 FID at 4 steps and 2.38 at 50 steps on ImageNet-256², with FID not degrading as budget grows from 4 to 50 NFE. The recipe transfers to text-to-image. Code and models are open-sourced on GitHub.

Original post →

More from Multimodal

Multimodal channel →