DC-SAE hits 32x compression with faster diffusion convergence, beating DC-AE by 54.9% gFID

DAGroup-PKU · hf · 2026-10-01

Peking University's DC-SAE tackles the tradeoff between aggressive tokenizer compression and slow diffusion convergence by pairing semantic encoders for high compression with a pixel-level encoder that preserves fine details.

On ImageNet 512×512 it achieves 32x spatial compression with 29.79 PSNR and 3.37 gFID, outperforming the previous best high-compression tokenizer DC-AE by 13.5% on PSNR and 54.9% on gFID, with comparable throughput and faster training convergence. A 1.6B DiT with DC-SAE scores 0.84 GenEval and 86.007 DPG-Bench for text-to-image at 1024×1024.

Original post →

More from Research

Research channel →