Anthropic Paper: Verbalizable Representations Form a Global Workspace in Language Models
rgblong · x · 2026-10-10
Anthropic published a new Transformer Circuits paper, "Verbalizable Representations Form a Global Workspace in Language Models," by Jack Lindsey, Wes Gurnee and other core members.
Key findings:
- LLMs appear to have developed a human-like functional split: a small privileged set of representations available for report, modulation, and flexible reasoning, sitting atop a much larger volume of automatic processing
- The team introduces a new interpretability technique that surfaces the concepts a model is poised to verbalize at any point, enabling measurement and intervention on the model's thought processes
- Commentators argue that even setting aside consciousness/GWT interpretations, it offers a compelling portrait of a rich internal world
The thread also points to the related jspace paper and accompanying commentary.
More from Research
- UPenn's GRASP lab releases OctoSense: 8.5TB multi-sensor robotics dataset with event cameras — RexDouglass · 2026-10-10
- Claim: OpenAI model cracks 'quantum-only' problems on a regular laptop — anirbanbandyo · 2026-10-10
- Mnemos: an engram-inspired memory architecture doubles short-term recall vs Claude's native memory — RileyRalmuto · 2026-10-10
- Startup builds 8,000 sq ft lab in 30 days to test AI-proposed molecules in ~72 hours — 141_1337 · 2026-10-10
- An LLM workflow mines multilingual citizen-science comments for species interaction observations — abenitezburraco · 2026-10-10
- Evolution isn't random search: MIT's Akarsh Kumar on ASAL and 262,144 artificial worlds — Machine Learning Street Talk · 2026-10-10