New Anthropic Research Can Read Claude's Internal Thoughts
AnthropicAI · x · 2026-07-07
Anthropic published a paper introducing an interpretability method called "J-space" (derived from the Jacobian matrix). Existing within the model's internal neural activations, J-space is distinct from Claude's output text or chain-of-thought, enabling the model to "think" about concepts without explicitly writing them out.
By observing J-space, researchers can read, audit, and shape what Claude is currently thinking. This provides tools to maintain trustworthiness as model capabilities improve, and also hints at striking similarities between language models and the human mind.
Related event: Anthropic Discovers Global Workspace Inside Claude(102 posts)→
More from Research
- Navier-Stokes, Riemann, P vs NP: what this week's math buzzwords mean for you — koltregaskes · 2026-09-11
- Fruit fly connectome LLM weights land on Hugging Face, transformers-compatible — ngxson · 2026-09-11
- Fruit fly brain as an LLM: connectome-driven language model demo goes live — ngxson · 2026-09-11
- Harry Collins: LLMs can't do frontier science because they can't invent new language — whoamisri · 2026-09-11
- The Waymo effect: how AI is quietly making research less collaborative — JohnHammersley · 2026-09-11
- Causal-only attention for non-generative tasks is wasteful, argues HF engineer — antoine_chaffin · 2026-09-11