V-RAE: New Paradigm for Video Generation via Foundation Model Representations

NUS and Oxford propose V-RAE, which builds a compact latent space on frozen vision foundation model representations with lightweight temporal pooling, replacing traditional VAE compression for video generation. Experiments on K600 reveal a reconstruction-generation quality gap while offering a new paradigm.

2026-08-25 ~ 2026-08-27 · 2 related posts