LDM-is-AE paper: DiT as autoencoder removes pretrained VAE, hits FID 1.80 on ImageNet
serrjoa · x · 2026-09-30
Latent Diffusion Models typically use a two-stage pipeline: a pretrained autoencoder (VAE) defines the latent space, then a diffusion model is trained inside it — creating a representation mismatch, since the latent space is optimized for reconstruction rather than denoising dynamics.
Key idea
- The paper observes that the LDM itself is an autoencoder: the DiT backbone performs a latent-to-feature-to-latent transformation at every denoising step, which can be interpreted as an internal decode–encode process.
- The authors split the DiT backbone into reciprocal components DiT-E (encoding) and DiT-D (decoding), imposing image-space supervision on intermediate features across all timesteps to form an explicit latent→image→latent path.
- This enables end-to-end, single-stage training that jointly learns latent representations and denoising, eliminating the need for a separately trained tokenizer (VAE).
Results: LDM-is-AE improves ImageNet generation to FID 1.80. Authors include Zhengqiang Zhang and Lei Zhang (arXiv:2609.37080).
More from Research
- ProDyGS Turns Monocular Video Into Pseudo-Multi-Views for Dynamic Gaussian Splatting — kwangmoo_yi · 2026-09-30
- From PDF archives to a trainable model: an open-source fine-tuning data workflow — Puzzleheaded_Box2842 · 2026-09-30
- Depth may be the next scaling axis — if you pick the right residual connections — FinanceYF5 · 2026-09-30
- DepthBench aims to settle the race to beat the 'depth curse' in deep LLMs — FinanceYF5 · 2026-09-30
- Researchers found a 'pain direction' in 25 models; someone claims to have weaponized it on a local model — ZeroStateReflex · 2026-09-30
- Training an AI agent on its own explanations improves coding—no teacher, no verifier, no RL — CatAstro_Piyush · 2026-09-30