Vision models may understand world structure better than LLMs
khademinori · x · 2026-08-15
François Fleuret argues that vision models understand the structure of the world projected on an image far better than LLMs understand the world projected into words. He suggests it is hard not to tap into vision models to improve abstract thinking in LLMs.
More from Research
- Why DeepSeek's dsh/cordis is a big deal: LH tasks and Harness meta tuning — EstablishmentOdd785 · 2026-08-15
- Stanford Virtual Embryo Challenge draws 401 researchers and 295 teams in first week — anshulkundaje · 2026-08-15
- TrackCraft3R: Repurposing Video Diffusion Transformers for Dense 3D Tracking — rsasaki0109 · 2026-08-15
- Zhejiang University open-sources Polaris: AI pipeline for full research workflow — 机器之心 · 2026-08-15
- Will Transformers dominate until the 2040s? Deep dive on architecture evolution — Concern-Excellent · 2026-08-15
- Data Pyramid framework boosts embodied manipulation skills — jiqizhixin · 2026-08-15