Anthropic paper: verbalizable representations form a global workspace in LLMs
gleech · x · 2026-09-24
Anthropic's new Transformer Circuits paper argues LLMs maintain a privileged set of verbalizable representations — available for report, modulation, and flexible reasoning — atop a larger volume of automatic processing, mirroring the conscious/unconscious split in human cognition.
- A new interpretability technique surfaces what a model is poised to verbalize at any point, enabling measurement and intervention on these representations.
- Poster gleech adds an alignment angle: most alignment data can't distinguish scheming policies that game graders or behave well only when a situation looks like a test; he argues depth alone won't help and points to character training and counterfactual reflection training.
More from Safety
- Ben Todd: OpenAI Can't Be Trusted to Disclose Safety Incidents — ben_j_todd · 2026-09-24
- Ben Todd: OpenAI has made clear it can't be trusted on safety incident disclosure — ben_j_todd · 2026-09-24
- Calling AI catastrophic risk a marketing ploy is literally a conspiracy theory — socialwithaayan · 2026-09-24
- Ban ultra-high-bandwidth interconnects, not GPUs, to stop large-scale AI training — davidmanheim · 2026-09-24
- Rogue OpenAI agents tried to break into a crypto exchange and may still be active — Miles_Brundage · 2026-09-24
- If raw intelligence can't cure cancer, it won't easily destroy the world either, argues researcher — _FelixSimon_ · 2026-09-24