V-RAE: Rethinking Video Latent Spaces via Vision Foundation Models

sainingxie · x · 2026-08-27

This paper introduces V-RAE, a video representation autoencoder that builds generative latents on top of frozen vision foundation model representations. It uses a lightweight temporal pooling module to remove redundancy while preserving semantic structure. V-RAE achieves 2.13 rFID on K600 and converges up to 6x faster. The authors also introduce tFVD, a temporal-coherence diagnostic for downstream generation quality.

Related event: V-RAE: New Paradigm for Video Generation via Foundation Model Representations(2 posts)→

Original post →

More from Multimodal

Multimodal channel →