Anthropic Studies a Global Workspace Inside LLMs
omarsar0 · x · 2026-07-21
This post highlights Anthropic’s paper “Verbalizable Representations Form a Global Workspace in Language Models.”
Main idea
Using a new interpretability method called the Jacobian lens, the authors identify internal representations that a model can put into words. They argue these verbalizable representations behave like a global workspace:
- they can be reported verbally,
- deliberately summoned and held in mind,
- used to carry intermediate reasoning steps,
- and broadcast more broadly through the model than other representations.
What the paper claims
- The workspace is bandwidth-limited and holds only a small set of concepts at once.
- It appears in an intermediate band of layers.
- It can expose signs of strategic deliberation, evaluation awareness, and other dispositions that may not appear in final outputs.
- Post-training appears to install an “Assistant’s point of view” in this workspace.
- The authors also propose counterfactual reflection training, training the model only on what it would say if interrupted and asked to reflect.
Why it matters
The paper is presented as a practical window into a model’s unspoken thinking and a mechanistic account of when verbalized reasoning is doing real work versus just narrating after the fact.
More from Research
- Project APE launches CRED to test whether LLMs can verify research errors — soumitrashukla9 · 2026-07-22
- Project APE finds verifier reliability drops when papers contain multiple errors — soumitrashukla9 · 2026-07-22
- Project APE says verifier costs fell about 90x in a year as Chinese open models lead — soumitrashukla9 · 2026-07-22
- OpenAI-linked paper says capability RL can make models more reward-seeking — MariusHobbhahn · 2026-07-22
- Project APE builds its verifier benchmark from 100 AI-written papers with injected errors — soumitrashukla9 · 2026-07-22
- Paper proposes a CRED taxonomy and benchmark to measure research-error detectors — soumitrashukla9 · 2026-07-22