Interpretability paper finds a faithful internal world model inside taxiGPT
__nmca__ · x · 2026-09-23
The new paper "World Modeling in Transformers" (with authors from MATS) asks whether AI models can recover world models purely from data.
In taxiGPT, the authors found a faithful internal map of the world, and traced the model's failures to how that map is stored in superposition. The finding directly counters the common "LLMs don't model the world" argument — empiricists looked inside and found a little model of the world.
Related event: Study Finds Faithful World Models Inside Transformers(2 posts)→
More from AGI Musings
- 'The New Anthropic Model Is Wonderful': Insider Says We're Far from the Pacing Event — iruletheworldmo · 2026-09-23
- Jensen Huang on AI safety: if labs can't contain their experiments, 'we have to shut the labs down' — ShakeelHashim · 2026-09-23
- Restaurant manager on AI booking calls: the fake conversationality feels condescending — annetgriffin · 2026-09-23
- Will AI kill folders? Predicting a 'one bucket' era where search does the organizing — NickPassig · 2026-09-23
- Al Gore: all AI data centers emit less than the world's uncovered landfills — jeffclune · 2026-09-23
- Uncle Bob: AI changes nothing—complexity, not tooling, still makes software slow — blaizedsouza · 2026-09-23