Oxford & NUS Propose V-RAE: Replacing Pixel Reconstruction with Pretrained Visual Representations for Video Generation

jiqizhixin · x · 2026-09-03

Oxford University and NUS present V-RAE (Video Representation Autoencoder), rethinking the foundation of video generation.

Original post →

More from Multimodal

Multimodal channel →