Anthropic interpretability roundup: Claude keeps a privileged global workspace it can report on
ctjlewis · x · 2026-09-28
A roundup of Anthropic's Transformer Circuits interpretability research. Key findings: Claude maintains a small privileged set of representations it can report on, control, and reason with atop automatic processing (a 'global workspace'); interference weights identified in a 1-layer transformer; emotion concept representations in Claude Sonnet 4.5 causally influence outputs; evidence of emergent introspective awareness in LLMs.
Other work includes training Claude to verbalize its own activations (Activation Oracles, Natural Language Autoencoders), the HeadVis attention visualization tool, and updates on sparse autoencoders and 'harm pressure'.
More from Research
- Hillock: open-source neuro-symbolic agent memory engine runs under 1.2GB VRAM — Equivalent-Flan-1590 · 2026-09-28
- PKU open-sources RayOrch, lineage-aware data-prep engine with up to 15.14x speedup — PekingUniversity · 2026-09-28
- YODAS v3 lands on Hugging Face: 1.1M hours, the largest open speech dataset ever — shinjiw_at_cmu · 2026-09-28
- Graph alignment is all you need: slides from CIRM workshop talk — marc_lelarge · 2026-09-28
- AI-picked catalyst dismissed by experts survived 1,000+ hours in acid — VraserX · 2026-09-28
- MICA: a 4.5MB Transformer-free LM splits into Ember and Flame, generation still weak — Silver_Employ2617 · 2026-09-28