DC-SAE hits 32x compression with faster diffusion convergence, beating DC-AE by 54.9% gFID
DAGroup-PKU · hf · 2026-10-01
Peking University's DC-SAE tackles the tradeoff between aggressive tokenizer compression and slow diffusion convergence by pairing semantic encoders for high compression with a pixel-level encoder that preserves fine details.
On ImageNet 512×512 it achieves 32x spatial compression with 29.79 PSNR and 3.37 gFID, outperforming the previous best high-compression tokenizer DC-AE by 13.5% on PSNR and 54.9% on gFID, with comparable throughput and faster training convergence. A 1.6B DiT with DC-SAE scores 0.84 GenEval and 86.007 DPG-Bench for text-to-image at 1024×1024.
More from Research
- ArchMap lands in Nature Genetics: code-free single-cell mapping onto reference atlases — burny_tech · 2026-10-01
- New Paper Argues Literary Tools Are Essential for Building Culturally Literate AI — begusgasper · 2026-10-01
- NVIDIA's Instant NuRec reconstructs a drivable 3DGS world from driving logs in ~1.5 seconds — rsasaki0109 · 2026-10-01
- KV-streams trains SWE agents 2x faster by preserving KV cache across compaction — burny_tech · 2026-10-01
- Researcher argues "ego" beats "persona" for describing LLM identity — repligate · 2026-10-01
- Silicon microring modulators push past 200Gb/s per lane to cut AI optical I/O power — jwt0625 · 2026-10-01