Paper: Transformers do have world models — failures trace to feature interference

burny_tech · x · 2026-09-23

A new arXiv paper, "World Modeling in Transformers," revisits TaxiGPT, a transformer trained on random Manhattan walks previously cited as lacking a coherent world model. Mechanistic analysis and causal interventions show it represents intersections and streets, tracks position, and uses a goal compass for navigation. Failures stem from interference between superposed intersection features that disrupt localization. "Affordance packing" (grouping intersections with identical legal moves) limits the damage, and the authors propose mechanistic indicators showing world-modeling capacities emerge at different training stages. The takeaway: shift from asking whether a model has a world model to mechanistically studying how it models the world.

Related event: Study Finds Faithful World Models Inside Transformers(2 posts)→

Original post →

More from Research

Research channel →