Karpathy says spatial skill is pure text prediction; researcher counters models train on images and GIS data

gleech · x · 2026-10-02

Karpathy emphasized that models' spatial/map abilities come "just from reading a ton of text and then predicting text — no images involved except how text responses are arranged in 2D."

gleech pushed back: such models are trained on images (and piles of GIS data), and he'd be surprised if newer ones haven't been trained in a GIS environment. The exchange spotlights a live debate over whether spatial reasoning is emergent from text or simply inherited from image/geospatial training data.

Related event: Karpathy Claims Models Learn Spatial Reasoning From Text Alone(2 posts)→

Original post →

More from Models

Models channel →