LDM-is-AE paper: DiT as autoencoder removes pretrained VAE, hits FID 1.80 on ImageNet

serrjoa · x · 2026-09-30

Latent Diffusion Models typically use a two-stage pipeline: a pretrained autoencoder (VAE) defines the latent space, then a diffusion model is trained inside it — creating a representation mismatch, since the latent space is optimized for reconstruction rather than denoising dynamics.

Key idea

Results: LDM-is-AE improves ImageNet generation to FID 1.80. Authors include Zhengqiang Zhang and Lei Zhang (arXiv:2609.37080).

Original post →

More from Research

Research channel →