Anthropic interpretability roundup: Claude keeps a privileged global workspace it can report on

ctjlewis · x · 2026-09-28

A roundup of Anthropic's Transformer Circuits interpretability research. Key findings: Claude maintains a small privileged set of representations it can report on, control, and reason with atop automatic processing (a 'global workspace'); interference weights identified in a 1-layer transformer; emotion concept representations in Claude Sonnet 4.5 causally influence outputs; evidence of emergent introspective awareness in LLMs.

Other work includes training Claude to verbalize its own activations (Activation Oracles, Natural Language Autoencoders), the HeadVis attention visualization tool, and updates on sparse autoencoders and 'harm pressure'.

Original post →

More from Research

Research channel →