Anthropic Studies a Global Workspace Inside LLMs

omarsar0 · x · 2026-07-21

This post highlights Anthropic’s paper “Verbalizable Representations Form a Global Workspace in Language Models.”

Main idea

Using a new interpretability method called the Jacobian lens, the authors identify internal representations that a model can put into words. They argue these verbalizable representations behave like a global workspace:

What the paper claims

Why it matters

The paper is presented as a practical window into a model’s unspoken thinking and a mechanistic account of when verbalized reasoning is doing real work versus just narrating after the fact.

Original post →

More from Research

Research channel →