VideoRAE Improves Latent Space for Video Generation

Zhihao Xie · hf · 2026-07-20

VideoRAE attempts to replace the traditional 3D-VAE with representations from a frozen video foundation model to serve as the latent space for video generation.

The core ideas are:

Experiments show that VideoRAE exhibits strong reconstruction quality and sets a new SOTA on UCF-101: AR and DiT generators achieve class-to-video gFVD scores of 40 and 93, respectively. Training convergence is roughly 5x faster compared to 3D-VAE baselines. In a 2B scale text-to-video experiment, replacing LTX-VAE with VideoRAE also led to faster convergence. Models and code will be open-sourced.

Original post →

More from Multimodal

Multimodal channel →