21 researchers release white paper on Visual General Intelligence as a path to AGI

HirokatuKataoka · x · 2026-10-10

A white paper on arXiv by Hirokatsu Kataoka and 20 co-authors (including Robert Geirhos, Deva Ramanan, Yilun Du, Jiajun Wu, Zhuang Liu, Andrew Davison) reconsiders intelligence from a vision-centered perspective. Just as GPT-style autoregressive modeling on web-scale text enabled transfer to unseen tasks, it asks what capabilities can emerge from visual modalities like images, video and geometry — and whether visual general intelligence (VGI) can be a pathway to AGI. Rather than offering a single definition, it maps out principles computer vision should pursue in the AGI era, visual input modalities, benchmarks, learning paradigms, and the relationship between vision and language when vision is taken as the core.

Original post →

More from Research

Research channel →