Video Generation Models Can Also Do Understanding

机器之心 · wechat · 2026-07-15

This article introduces a new Google DeepMind paper, 《Video Generation Models are General-Purpose Vision Learners》. The core method is called GenCeption, which transforms pre-trained video generation models into a general-purpose video understanding system controlled by text instructions.

Key Concept

Unified Multi-Task Approach

Data & Experiments

Generalization & Conclusion

Related event: DeepMind's GenCeption: Video Generation Models as General-Purpose Vision Learners(6 posts)→

Original post →

More from Multimodal

Multimodal channel →