DeepMind, Harvard and Stanford paper: visual world models may be a path to AGI
rohanpaul_ai · x · 2026-09-12
A paper from Google DeepMind with Harvard, Stanford and other top labs argues that visual AI could be a path to AGI: systems that build world models, remember changes, predict outcomes, and act. Key points:
- Most multimodal AI treats vision merely as input to a language model; the paper wants vision to do more of the thinking itself
- A capable visual system should learn directly from images, video, 3D structure, and interaction — understanding what exists, what changed, what is hidden, what might happen next, and what to look at before acting
- Video generation, reconstruction, persistent memory, continual learning, multimodal sensing, and robotics could be pieces of the same system
- Stop judging visual AI mainly by image Q&A, captions, or realistic video
Related event: DeepMind, Harvard and Stanford Say Visual AI May Be the Path to AGI(2 posts)→
More from AGI Musings
- Cambridge AI researcher David Krueger explains the "gradual disempowerment" doomsday scenario — ronbodkin · 2026-09-12
- Human societies thrive despite unaligned individuals — a fresh angle on AI alignment — ethanniser · 2026-09-12
- Answer: prosperous societies run on well-constructed institutions, not universal alignment — ethanniser · 2026-09-12
- 24 Fields Medalists sign declaration warning of "Severe Misalignment" of AI in mathematics — stevenstrogatz · 2026-09-12
- Christof Koch: simulating consciousness convincingly doesn't mean the system actually feels anything — pwlot · 2026-09-12
- Anthropic safety researcher Joe Benton quits to join METR, citing extinction-level AI risk — JacquesThibs · 2026-09-12