Interpretability paper finds a faithful internal world model inside taxiGPT

__nmca__ · x · 2026-09-23

The new paper "World Modeling in Transformers" (with authors from MATS) asks whether AI models can recover world models purely from data.

In taxiGPT, the authors found a faithful internal map of the world, and traced the model's failures to how that map is stored in superposition. The finding directly counters the common "LLMs don't model the world" argument — empiricists looked inside and found a little model of the world.

Related event: Study Finds Faithful World Models Inside Transformers(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →