Karpathy: it's all text prediction — no images involved, only 2D arrangement of responses

karpathy · x · 2026-10-02

Karpathy emphasizes that these model capabilities come purely from reading a ton of text and predicting text. No images are involved, except in how text responses are arranged in 2D.

Related event: Karpathy Claims Models Learn Spatial Reasoning From Text Alone(2 posts)→

Original post →

More from Models

Models channel →