HiRAE cuts ImageNet-256 reconstruction FID to 0.209 with residual-budget autoencoding
Xuanyu Zhu · hf · 2026-09-30
HiRAE is a Hierarchical Representation Autoencoder that improves reconstruction fidelity of pretrained visual representations for generative modeling:
- Groups encoder layers by depth and learns residual corrections to the deepest representation, with group-wise norm caps bounding corrections (tighter budgets for shallower groups).
- HiRAE-24 keeps the same latent token count and channel dimension.
Results: on ImageNet-256, reconstruction FID drops from 0.299 to 0.209 vs. RAEv2 while maintaining competitive guided generation quality. For text-to-image, GenEval improves from 84.86 to 87.70 post-fine-tuning, with gains on DPG-Bench and GenAI-Bench as well.
More from Multimodal
- Four imaginary tokens for Midjourney v8.2 produce memory ghosts and bone echoes — LudovicCreator · 2026-09-30
- Hyper-personalized music is BS: music is culture and inherently social, argues developer — jordiponsdotme · 2026-09-30
- NUS Proposes StoryEngine: A State-Grounded Agentic Framework for Coherent Long-Form Video Storytelling — NationalUniversityofSingapore · 2026-09-30
- One Year of Local Image Generation: Why Civitai and ComfyUI Both Fall Short — BenDLH · 2026-09-30
- Opus made a launch video for Violetto 1B in 50 minutes amid zero media coverage — tensorqt · 2026-09-30
- Meshy hits $100M ARR in under two years as GPT-6 Astra stirs the AI 3D debate — 量子位 · 2026-09-30