New Paper: Transformers Hold Faithful Internal World Maps but Fail to Act on Them

xuanalogue · x · 2026-09-23

A new paper, World Modeling in Transformers, asks whether AI models can recover world models purely from data. Using taxiGPT as the subject, the authors find a faithful internal map—yet the model fails to execute it behaviorally.

Mechanistic interpretability traces the failures to how the map is stored in superposition. Reposters highlight that NextLat has the best world model mechanistically, though still imperfect.

Related event: Paper: Transformers Do Have Faithful World Models, Just Scrambled(3 posts)→

Original post →

More from Models

Models channel →