Applying a Cipher to the Corpus Suffices to Dismiss Theories of Meaningful Internal States in LLMs
gerardsans · x · 2026-10-08
Gerard Sans continues his critique of LLM internal-state theories, connected to the World Properties without World Models preprint discussion.
Argument:
- Text or pixels capture token co-occurrence in data; the data carries structures an observer recognises and projects meaning onto
- Apply a cipher unknown to the observer and the projection of meaning vanishes; the same applies to someone who doesn't understand English or lacks world knowledge — you can only see what you recognise
- Yet the sampler (model) works identically on the original or obfuscated corpus
The author concludes this suffices to dismiss any theory of internal states holding meaning beyond co-occurrence distributions.
Related event: Word Vectors Without Transformers Question LLM World Models(2 posts)→
More from AGI Musings
- We used to celebrate the technologist — now AI has us celebrating the model — martyjbeard · 2026-10-08
- OpenAI releases 372 math results in one day, including an AI proof of the Unique Games Conjecture — LukeW · 2026-10-08
- LM4Sci Workshop at COLM'26 to Gather LLM-for-Science Researchers Across Anthropic, Princeton, UW — ChengleiSi · 2026-10-08
- Researcher says agent token spend tops his rent, courts frontier labs for access — iScienceLuvr · 2026-10-08
- Stanford Professor: Stop Spinning Negative Results, Brutally Cull Ideas That Miss Milestones — anshulkundaje · 2026-10-08
- OSHBuilt's bet: MCP servers become the standard interface that lets AI agents source US manufacturing — hudzah · 2026-10-08