Paper: Transformers do have world models — failures trace to feature interference
burny_tech · x · 2026-09-23
A new arXiv paper, "World Modeling in Transformers," revisits TaxiGPT, a transformer trained on random Manhattan walks previously cited as lacking a coherent world model. Mechanistic analysis and causal interventions show it represents intersections and streets, tracks position, and uses a goal compass for navigation. Failures stem from interference between superposed intersection features that disrupt localization. "Affordance packing" (grouping intersections with identical legal moves) limits the damage, and the authors propose mechanistic indicators showing world-modeling capacities emerge at different training stages. The takeaway: shift from asking whether a model has a world model to mechanistically studying how it models the world.
Related event: Study Finds Faithful World Models Inside Transformers(2 posts)→
More from Research
- AI models now match superforecasters on ForecastBench; rematch set for October — burny_tech · 2026-09-23
- Dev uses Opus 5.5 with Lean to formally verify Claude Agent SDK, yielding 16 bug-fix PRs — bcherny · 2026-09-23
- Mathematicians, not just LLMs, made AI's math breakthroughs possible, scholars argue — tak3sh8 · 2026-09-23
- AI-enabled drug discovery cuts discovery time by 15-80%, McKinsey research finds — menhguin · 2026-09-23
- Gemini training details dissected: groupwise reward redistribution to fight reward hacking — nrehiew_ · 2026-09-23
- New model's architecture is 'vanilla': SWA plus MoE with no shared experts, unlike DeepSeek — nrehiew_ · 2026-09-23