Oxford & NUS Propose V-RAE: Replacing Pixel Reconstruction with Pretrained Visual Representations for Video Generation
jiqizhixin · x · 2026-09-03
Oxford University and NUS present V-RAE (Video Representation Autoencoder), rethinking the foundation of video generation.
- Every video generator compresses video into a latent space first, but traditional VAEs optimize pixel reconstruction — faithfully reconstructing real video differs from easily generating new video.
- V-RAE skips pixel-reconstruction objectives entirely, using the high-level representations learned by a pretrained vision model as the generation space, extending the RAE idea from images to video.
- The diffusion model thus learns from a latent space with rich semantic structure rather than one optimized for texture fidelity.
- Results on Kinetics-600 are shown (the post is truncated; see original for numbers).
More from Multimodal
- Runway launches Dev MCP server for coding agents — tlakomy · 2026-09-03
- MiniMax H3 open weights one month in: developer builds generative video classroom with real-time animated explanations — alejandroll10 · 2026-09-03
- ComfyUI to announce MiniMax H3 sync sound challenge winners in special livestream — MiniMax_AI · 2026-09-03
- Wan 2.2 Image-to-Video Tested in ComfyUI With Two-Stage High-to-Low Noise KSampler Pass — VictorVisuals · 2026-09-03
- AI influencers reshape content creation as Seedance 2.5 targets influencer vlogs — aftahi_ai · 2026-09-03
- Seedance 2.5 tested for AI influencer vlogs: 30-second clips with consistent identity — aftahi_ai · 2026-09-03