HiRAE cuts ImageNet-256 reconstruction FID to 0.209 with residual-budget autoencoding

Xuanyu Zhu · hf · 2026-09-30

HiRAE is a Hierarchical Representation Autoencoder that improves reconstruction fidelity of pretrained visual representations for generative modeling:

Results: on ImageNet-256, reconstruction FID drops from 0.299 to 0.209 vs. RAEv2 while maintaining competitive guided generation quality. For text-to-image, GenEval improves from 84.86 to 87.70 post-fine-tuning, with gains on DPG-Bench and GenAI-Bench as well.

Original post →

More from Multimodal

Multimodal channel →