DeepMind, Harvard, Stanford Argue Visual Intelligence May Be a Path to AGI
rohanpaul_ai · x · 2026-09-12
A white paper by 21 researchers from Google DeepMind, Harvard, Stanford, Oxford and other labs argues that vision-centered AI could offer a path to AGI, termed "Visual General Intelligence" (VGI).
- The analogy: in language, the GPT series showed transfer to unseen tasks via Transformer architecture, autoregressive training on web-scale text, and aggressive scaling. The paper asks what capabilities could emerge from images, video, and geometry.
- Critique of the status quo: most multimodal AI treats vision as input to a language model; the paper wants vision to do more of the thinking itself — learning directly from visual data to build world models, remember changes, predict outcomes, and act.
- Rather than a single definition, the paper maps the principles CV should pursue in the AGI era, visual modalities, benchmarks, learning paradigms, and vision's relationship with language and other modalities.
Related event: DeepMind, Harvard and Stanford Say Visual AI May Be the Path to AGI(2 posts)→
More from Research
- Persistent-memory denoiser hits 99.9% on Sudoku-Extreme with under 250k parameters — LucaAmb · 2026-09-12
- MIT's Astra invents its own scientific instruments, finds metamaterial 2.1x stronger than reference — ProfBuehlerMIT · 2026-09-12
- GPT-6 fails to improve on molecular property prediction, fueling AGI skepticism — GaryMarcus · 2026-09-12
- Fruit fly brain connectome simulated to play Beat Saber, startling researchers — LinusEkenstam · 2026-09-12
- Fruit fly connectome plays chess at ~700 Elo in playable online demo — dejavucoder · 2026-09-12
- WUJI open-sources MINT: 3D hand trajectory and camera pose reconstruction from egocentric video — chris_j_paxton · 2026-09-12