DeepMind says video generators can become general-purpose vision models
jiqizhixin · x · 2026-07-23
Google DeepMind and collaborators introduce GenCeption, arguing that video generation models can serve as a shared backbone for general-purpose vision.
Main idea
- Instead of training separate models for each perception task, the system starts from a pretrained text-to-video generator.
- That backbone learns spatiotemporal structure and language alignment, then transfers to downstream vision tasks.
Reported results
- It matches or beats specialist models such as DepthAnything, SAM, and VGGT on depth estimation, surface normals, camera pose, expression-referenced segmentation, and 3D keypoint prediction.
- It also outperforms pretraining baselines like V-JEPA and Video MAE.
- The authors say it uses 7x to 500x less training data.
- A model trained only on synthetic human videos still generalizes to real-world footage, including animals and robots.
The paper's claim is broad: video generation models may be a general-purpose vision learner, not just a media generator.
More from Multimodal
- Taluna.ai shows a stylized animation clip in “Legend of Sei” — ScheduleNo9532 · 2026-07-23
- Seedance2 video-to-video workflow uses depth video plus character references — Horror_Dirt6176 · 2026-07-23
- Wired says Jibo’s successor is an AI wearable that turns family life into slop — Wired AI · 2026-07-23
- Blomkamp’s 13-minute sci-fi short was made entirely with Dreamina Seedance 2.0 4K — testingcatalog · 2026-07-23
- AI-generated World Cup final clips hit 28 million views — techhalla · 2026-07-23
- Midjourney SREF code 3203007452 targets a cute minimalism look — aziz4ai · 2026-07-23