VideoRAE Improves Latent Space for Video Generation
Zhihao Xie · hf · 2026-07-20
VideoRAE attempts to replace the traditional 3D-VAE with representations from a frozen video foundation model to serve as the latent space for video generation.
The core ideas are:
- Extracting multi-scale hierarchical features from a frozen video foundation model;
- Compressing them into generation-friendly latents using a lightweight 1D self-attention projector;
- Supporting both continuous latents for Diffusion Transformers and discrete tokens for Autoregressive models.
Experiments show that VideoRAE exhibits strong reconstruction quality and sets a new SOTA on UCF-101: AR and DiT generators achieve class-to-video gFVD scores of 40 and 93, respectively. Training convergence is roughly 5x faster compared to 3D-VAE baselines. In a 2B scale text-to-video experiment, replacing LTX-VAE with VideoRAE also led to faster convergence. Models and code will be open-sourced.
More from Multimodal
- Seedance 2.0 demo turns ketchup on spaghetti in Rome into an AI reaction meme — azed_ai · 2026-07-21
- A reusable “Lunar Eclipse Dreamscape” prompt comes with multiple example renders — LudovicCreator · 2026-07-21
- Midjourney 8.2 preview shows a double-exposure prompt with strong style control — michaelrabone · 2026-07-21
- Travel MCP Server adds flight, hotel, weather and budget tools for agents — modelcontextprotocol · 2026-07-21
- Douyin Video Analysis MCP turns share links into structured video summaries — modelcontextprotocol · 2026-07-21
- Synthesia launches Dubbing 2.0 with 130+ languages and lip-sync video translation — synthesiaIO · 2026-07-21