V-RAE: Rethinking Video Latent Spaces via Vision Foundation Models
sainingxie · x · 2026-08-27
This paper introduces V-RAE, a video representation autoencoder that builds generative latents on top of frozen vision foundation model representations. It uses a lightweight temporal pooling module to remove redundancy while preserving semantic structure. V-RAE achieves 2.13 rFID on K600 and converges up to 6x faster. The authors also introduce tFVD, a temporal-coherence diagnostic for downstream generation quality.
More from Multimodal
- Nvidia Releases 4-Step Versions of Cosmos3 Super T2I and I2V Models — q5sys · 2026-08-27
- ComfyUI and MiniMax Launch H3 Sync Sound Challenge — Comfy-Org · 2026-08-27
- Orpheus Concept Short Created with Turbo LoRA, Workflow Open-Sourced — EasternAd8821 · 2026-08-27
- Convert Any Stereo Track to Spatial Audio with Logic Pro — BLCNYY · 2026-08-27
- Original AI Music Video Created with MiniMax H3 R2V and Suno — HowToSD · 2026-08-27
- FixAnything: Refining 3D Renders into Photorealistic Videos via Video Generative Priors — orlitany · 2026-08-27