UniSpace: Unified Visual Representation without VAE

meituan-longcat · hf · 2026-08-24

UniSpace introduces a reparameterized pretrained vision transformer that unifies semantic understanding, high-fidelity reconstruction, and image generation within a single visual space, eliminating the need for a separate VAE.

Original post →

More from Multimodal

Multimodal channel →