Anthropic's Interpretability Research: Models Form a Global Workspace

aryaman2020 · x · 2026-08-08

The thread reviews Anthropic's series of interpretability research on Transformer Circuits Thread. It highlights the 'J Space paper', noting that the middle layers of language models form a global workspace with a privileged set of representations that can be reported on and controlled.

Additionally, the summary covers cutting-edge explorations including natural language autoencoders that translate internal states into text, the causal influence of emotion concepts on model outputs, and evidence of emergent introspective awareness in models regarding their own states.

Original post →

More from Research

Research channel →