DeepMind paper: Veo 3 shows emergent zero-shot reasoning, video models may become vision foundation models

RexDouglass · x · 2026-09-17

A Google DeepMind paper, "Video models are zero-shot learners and reasoners," shows Veo 3 zero-shot solves many untrained tasks: segmentation, edge detection, image editing, physical property understanding, tool-use simulation, and early visual reasoning like maze and symmetry solving — suggesting video models are on a path to unified vision foundation models, like LLMs for language.

Original post →

More from AGI Musings

AGI Musings channel →