V-RAE Rethinks Video Latent Spaces for Better Generation and Prediction

机器之心 · wechat · 2026-08-25

Researchers from the National University of Singapore and the University of Oxford propose V-RAE (Video Representation Autoencoder), rethinking the latent spaces used for video generation. Instead of learning latents based on pixel reconstruction like traditional VAEs, V-RAE uses frozen visual foundation models as encoders and employs lightweight temporal pooling to compress video features directly into semantic representations for generation and prediction.

Key Findings:

New Metric: tFVD

The paper introduces tFVD (Temporal FVD) to evaluate the temporal smoothness and prediction robustness of a latent space by interpolating between adjacent latent trajectories. tFVD shows a much stronger correlation with generation quality (0.62–0.92) than reconstruction metrics, indicating that a generative latent space must be "generatable" and stable, not just reconstructable.

Furthermore, V-RAE demonstrates significant advantages in future video prediction tasks, suggesting that this paradigm helps models learn the transitions of visual states more effectively.

Related event: V-RAE: New Paradigm for Video Generation via Foundation Model Representations(2 posts)→

Original post →

More from Research

Research channel →