21 researchers release white paper on Visual General Intelligence as a path to AGI
HirokatuKataoka · x · 2026-10-10
A white paper on arXiv by Hirokatsu Kataoka and 20 co-authors (including Robert Geirhos, Deva Ramanan, Yilun Du, Jiajun Wu, Zhuang Liu, Andrew Davison) reconsiders intelligence from a vision-centered perspective. Just as GPT-style autoregressive modeling on web-scale text enabled transfer to unseen tasks, it asks what capabilities can emerge from visual modalities like images, video and geometry — and whether visual general intelligence (VGI) can be a pathway to AGI. Rather than offering a single definition, it maps out principles computer vision should pursue in the AGI era, visual input modalities, benchmarks, learning paradigms, and the relationship between vision and language when vision is taken as the core.
More from Research
- AI-text detector misses 79.8% of abstracts rewritten by Meta's Muse-Glimmer, finds study — rohanpaul_ai · 2026-10-10
- Retrieval-centric deep learning: replacing weight matrices with vector databases — RobertTLange · 2026-10-10
- SJTU's UVTA teaches robot dexterous manipulation from 1,000 human tactile demos per task — siyuanhuang95 · 2026-10-10
- 20-minute explainer breaks down how Tesla trains FSD: 8 cameras, 36fps, 2B signals to 2 outputs — PTrubey · 2026-10-10
- Epoch AI: Over half of arXiv math papers in 3 subfields now acknowledge AI use — burny_tech · 2026-10-10
- Researcher faults standard Transformer diagrams for hiding self-attention as a black box — PlisSergey · 2026-10-10