DC-SAE: a dual-stream semantic-plus-pixel autoencoder cuts gFID 55% vs DC-AE at 32x compression

DC-SAE: Deep Compression Semantic Autoencoder for Faster Diffusion Convergence

Xu Huang, Ye Huang, Zijun Liao, Yuwei Niu, Xiaojie Li, Menghan Zhou, De Wen Soh, Xiaotong Li, Daquan Zhou

cs.CV

2026-09-30

A frozen semantic encoder plus an unconstrained pixel branch gives DC-SAE 32x compression: 29.79 PSNR and 3.37 gFID on ImageNet 512, beating DC-AE by 54.9% at 4.4x throughput.

What problem this solves

For latent diffusion models, the length of the latent sequence dominates compute. Deep-compression tokenizers such as DC-AE push spatial compression to 32x, which makes each forward pass cheaper but the latent harder to learn: diffusion training needs more steps to converge, and the saving gets eaten by the extra steps. The other line of work, representation autoencoders (RAE), swaps the VAE encoder for a pretrained semantic encoder like DINOv2. Those latents are structured and converge fast, but they top out around 16x compression and lose the texture, color statistics, and high-frequency detail that faithful reconstruction needs.

Pushing a semantic encoder to 32x directly does not work. Resizing the image to half resolution before DINOv2 yields roughly 15 PSNR at 256x256, and after 80 epochs of DiT training the model only reaches 7.15 gFID, so the fast-convergence property is gone too. MLLM-style post-hoc token merging (average pooling, learnable mergers) is worse. The diagnosis: semantic encoders are simply not trained for invertible reconstruction, and the information loss gets amplified at 32x.

Method

DC-SAE splits the conflict across two encoder streams, four parts in total:

The design is notable for what it skips: no KL regularization, no REPA-style representation alignment, no channel masking. The semantic stream gives the generator a learnable structure; the pixel stream gives reconstruction capacity. An ablation shows the two streams must be trained jointly inside one autoencoder. Late-Concat, which trains the pixel autoencoder alone and concatenates frozen DINOv2 features only at DiT training time, actually reconstructs better (31.91 vs 29.79 PSNR) but lands at 5.64 gFID after 80 epochs versus 4.02 for joint training. The gain comes from a jointly learned latent, not from merely exposing DiT to semantic features.

Results

ImageNet 512x512, both at 32x compression with a DiT-XL backbone:

MethodPSNR↑rFID↓gFID↓
DC-AE f32c3226.250.207.47
DC-SAE29.790.163.37

That is +13.5% PSNR and -54.9% gFID. Autoencoder throughput at 1024x1024 on an H200 (batch 32) is 76.8 vs 17.4 imgs/s (4.41x); at 256 and 512 the speedups are 4.64x and 4.75x.

Convergence ablation at 32x after 80 epochs: joint latent 4.27 gFID, semantic-only 7.15, pixel-only 11.04. On ImageNet 256x256, DC-SAE (f32c64) with DiT-XL plus a DDT wide head reaches 3.31 gFID and 180.71 IS; the same backbone on DC-AE gets 10.18, and DC-AE with the specialized USiT-H backbone still needs 3.89. At 16x, DC-SAE scores 27.80 PSNR and 3.09 gFID where RAE sits around 19 PSNR and 7.90 gFID.

For text-to-image, a 1.6B DiT with a Qwen3-1.7B text encoder at 1024x1024 with CFG scores 0.84 on GenEval and 86.007 on DPG-Bench. The 12B DC-Gen-FLUX.1-Krea baseline built on DC-AE scores 0.72 and 87.073.

The most honest experiment in the paper: masking the semantic branch entirely barely changes reconstruction, while masking the pixel branch destroys it. Reconstruction lives in the pixel stream; the semantic stream contributes a structured latent space for the generator.

Why it matters

For anyone training high-resolution image or video generators, tokenizer compression multiplies directly into sequence length. DC-SAE shows that at 32x you do not have to trade reconstruction fidelity against convergence speed, and you get there without stacking latent-space regularizers. Code and weights are already public from the PKU group. The pre-merge finding is also useful for MLLM vision-token design: compress before semantic abstraction, not after.

Limitations

The 54.9% gFID improvement holds under same-backbone comparison. In the same table, the 8x-compression EDM2-XXL reaches 1.91 gFID and DC-AE with a USiT-2B backbone reaches 2.90, both below DC-SAE's 3.37. The text-to-image comparison is not apples-to-apples: the baseline is a distilled 12B model trained on different data (the authors use internal data plus LAION-COCO), and DPG-Bench is slightly ahead for the baseline by about one point. That experiment shows a competitive 1.6B model, not a tokenizer verdict; the paper flags the configuration mismatch itself.

Evidence for faster convergence is gFID-per-epoch curves plus autoencoder throughput; an end-to-end wall-clock cost match at equal gFID is not reported. Why the semantic branch accelerates DiT convergence remains a hypothesis after the masking experiment, with no mechanistic validation. And the paper covers images only, no video yet.

Terms

Source

Related papers

All paper explainers