DeepMind researcher: all model inputs are arbitrary binary codes, so images aren't magic for understanding
AndrewLampinen · x · 2026-09-28
In an X exchange with Stanford linguist Chris Potts, DeepMind researcher Andrew Lampinen argues that all signals these systems process are arbitrary binary codes—we merely interpret chunks as different modalities. Since much of what an image or sensor reading conveys could be said in a chat, he questions the assumption that non-language inputs carry a magical "meaning-ness" essential to understanding, pushing back on the popular claim that multimodal grounding is required for genuine comprehension.
More from Research
- Study of 2,170 GitHub projects maps how the fast, low-cost Jev decision model is used in the wild — CUHK-CSE · 2026-09-28
- FoMo uses diffusion trajectory forking moments as annotation-free perceptual distance labels for IQA — SeoulNatlUniv · 2026-09-28
- CARD combines cluster-level LoRA adapters and reward-guided decoding for scalable LLM personalization — Yutong Song · 2026-09-28
- PsPLUG: explicit style instructions cause 'personalization collapse' in LLMs, lightweight plug-in fixes it — Yutong Song · 2026-09-28
- CMU's TrackEverything tracks all visible points in 1000+ frame videos within 40GB GPU memory — CarnegieMellonU · 2026-09-28
- ZooWork-ShopRanker: open 0.6B-8B e-commerce rerankers aligned to LLM-judged shopping preferences — _reachsumit · 2026-09-28