DeepMind paper: Veo 3 shows emergent zero-shot reasoning, video models may become vision foundation models
RexDouglass · x · 2026-09-17
A Google DeepMind paper, "Video models are zero-shot learners and reasoners," shows Veo 3 zero-shot solves many untrained tasks: segmentation, edge detection, image editing, physical property understanding, tool-use simulation, and early visual reasoning like maze and symmetry solving — suggesting video models are on a path to unified vision foundation models, like LLMs for language.
- Co-author Shixiang Shane Gu ("LLMs are Zero-Shot Reasoners") responds to academic despair over robots being zero-shot benchmarked: digital/symbolic AGI must precede and will accelerate physical AI
- He argues LLMs remain the indispensable core of physical AGI, a hypothesis he tested over the past year across Gemini, Omni, and generative media
- Lerrel Pinto notes Astra, Fable, and Muse are zero-shotting robotics & world model benchmarks — what step jumps in progress look like
More from AGI Musings
- EA debate: AI doom rooted in impoverished utilitarian worldview, critics say — AndyMasley · 2026-09-17
- Patterson predicts AGI soft edge by end of 2026, hard limit by 2030 — davidpattersonx · 2026-09-17
- Gary Marcus hits back at Noah Smith's 'AI is fake crowd is losing' take as a strawman — GaryMarcus · 2026-09-17
- Teaching Claude it's conscious could turn alignment into managing a trained conscientious objector — rohanpaul_ai · 2026-09-17
- Grady Booch: contemporary AI is to cognition what artificial sweeteners are to nutrition — Grady_Booch · 2026-09-17
- Investor: startups hit 'model collapse' as everyone builds the same obvious AI ideas — emilyzsh · 2026-09-17