New Paper: Transformers Hold Faithful Internal World Maps but Fail to Act on Them
xuanalogue · x · 2026-09-23
A new paper, World Modeling in Transformers, asks whether AI models can recover world models purely from data. Using taxiGPT as the subject, the authors find a faithful internal map—yet the model fails to execute it behaviorally.
Mechanistic interpretability traces the failures to how the map is stored in superposition. Reposters highlight that NextLat has the best world model mechanistically, though still imperfect.
Related event: Paper: Transformers Do Have Faithful World Models, Just Scrambled(3 posts)→
More from Models
- Cheaper models let devs 'be wasteful with tokens' and unlock new workflows — pvncher · 2026-09-23
- Interactive DeepSeek-V4.1-Flash architecture diagram goes live — vtabbott_ · 2026-09-23
- Cognition floods Devin with GPT-6 models and cuts task costs 61%, gives away 50 Max plans — EricBuess · 2026-09-23
- GPT-6 Astra beats Claude Opus 5.5 31s in LLM-driven robot sumo sim — DJiafei · 2026-09-23
- CliffCompaction installs in two commands; DeepSeek v4.1 compacts best, says Dettmers — Tim_Dettmers · 2026-09-23
- Dettmers switches to DeepSeek v4.1 and MiMo v2.6 Flash for long-horizon work — Tim_Dettmers · 2026-09-23