Study Finds Faithful World Models Inside Transformers
An arXiv paper on taxiGPT, trained on Manhattan random-walk data, found a faithful internal world map, challenging the claim that Transformers lack world models. Feature superposition interferes with, but does not eliminate, the learned world model.
2026-09-23 ~ 2026-09-23 · 2 related posts
- Interpretability paper finds a faithful internal world model inside taxiGPT — __nmca__ · 2026-09-23
- Paper: Transformers do have world models — failures trace to feature interference — burny_tech · 2026-09-23