New Anthropic Research Can Read Claude's Internal Thoughts

AnthropicAI · x · 2026-07-07

Anthropic published a paper introducing an interpretability method called "J-space" (derived from the Jacobian matrix). Existing within the model's internal neural activations, J-space is distinct from Claude's output text or chain-of-thought, enabling the model to "think" about concepts without explicitly writing them out.

By observing J-space, researchers can read, audit, and shape what Claude is currently thinking. This provides tools to maintain trustworthiness as model capabilities improve, and also hints at striking similarities between language models and the human mind.

Related event: Anthropic Discovers Global Workspace Inside Claude(102 posts)→

Original post →

More from Research

Research channel →