Anthropic interpretability researcher tells Tim O'Reilly: an LLM's world model is readable

mlpowered · x · 2026-09-17

Tim O'Reilly interviewed Emmanuel Ameisen, a researcher on Anthropic's AI interpretability team, going deep on what happens inside an LLM while it processes text. Core claims: prediction demands a world model; that world model is readable; and it is at work in every token. Researchers study activation patterns between layers to see which patterns correspond to particular ideas, and can even intervene by replacing activations with different values to observe how behavior changes. O'Reilly notes the most provocative angle is what studying LLMs might teach us about being human.

Original post →

More from Research

Research channel →