Anthropic interpretability researcher tells Tim O'Reilly: an LLM's world model is readable
mlpowered · x · 2026-09-17
Tim O'Reilly interviewed Emmanuel Ameisen, a researcher on Anthropic's AI interpretability team, going deep on what happens inside an LLM while it processes text. Core claims: prediction demands a world model; that world model is readable; and it is at work in every token. Researchers study activation patterns between layers to see which patterns correspond to particular ideas, and can even intervene by replacing activations with different values to observe how behavior changes. O'Reilly notes the most provocative angle is what studying LLMs might teach us about being human.
More from Research
- Boltz fused kernels deliver nearly 10x protein folding throughput gain — AllThingsApx · 2026-09-17
- TMLR desk rejects all 10 suspect papers after author verification calls; 3 couldn't answer basic questions — RexDouglass · 2026-09-17
- VoiceMem: Shuicheng Yan's team builds a dual-brain memory system that remembers tone — jiqizhixin · 2026-09-17
- Common Crawl puts crawl archives on Hugging Face Storage Bucket, with a getting-started guide — vanstriendaniel · 2026-09-17
- TMLR editor interviewed authors of low-quality submissions: they had no idea what their papers said — TuhinChakr · 2026-09-17
- Schmidhuber: I published the first concrete RSI algorithms in 1987, now RSI is driving AI's future — SchmidhuberAI · 2026-09-17