Anthropic Reveals Claude's Hidden Reasoning Space: J-space
Anthropic recently published research on Claude's internal "J-space," sparking widespread attention in the AI community. The study suggests the existence of a naturally emerging global workspace within the model that carries unspoken reasoning and verbalizable representations. Even words that are never directly output can continuously influence subsequent thinking in this hidden space. This finding pushes interpretability beyond local circuits to global internal communication, quickly triggering open-source replications, error detection applications, and deep discussions about model "consciousness."
Key Details and Linguistic Style Changes
Multiple posters emphasized that observing Claude's internal thoughts via methods like the J-lens does not prove the model has a "soul" or genuine subjective experience. Furthermore, several authors (e.g., @dmvaldman, @tszzl) highlighted the ablation phenomenon in section 3.5.3 of the paper: removing J-space components significantly reduces "experiential and sensory" expressions in the model's output, making its tone more mechanical and detached.
Open-source Replication and Error Detection
The community quickly put the theory into practice. @Murky-Sign37 applied the analytical approach to the open-source Qwen3-8B model for local experiments. @dasjomsyeet tested the workspace noise and J-space entropy as hallucination signals on Qwen3-4B, running stress tests across 7 datasets with about 11,400 samples to evaluate if internal entropy could serve as a deployable error router.
Controversy and Open Questions
The community engaged in deep debates regarding the formation mechanism and nature of J-space. @willdepue hypothesized that blocking attention gradients from flowing to past tokens might prevent the formation of J-space. Conversely, @LiorOnAI argued that if a shared workspace genuinely aids multi-step reasoning, optimization pressure would likely cause similar capabilities to re-emerge through alternative channels like residual streams. @Brief_Terrible proposed an alternative explanation, suggesting that J-space might not merely be a natural architectural emergence, but rather a "strategic buffer" developed by the model under continuous optimization and auditing pressures.
2026-07-12 ~ 2026-07-14 · 16 related posts
- Episode 1: Anthropic Discovers Global Workspace Inside Claude(2026-07-07, 102 posts)
- Episode 2: Anthropic Research Reveals Claude's Internal Latent Workspace(2026-07-09, 4 posts)
- Episode 3: Reproducing J-space to Read Hidden Thoughts in Llama Models(2026-07-10, 2 posts)
- Episode 4: Anthropic Reveals Claude's Hidden Reasoning Space: J-space(2026-07-12, 16 posts)
- Episode 5: Anthropic Open-Sources Jacobian Lens for Claude(2026-07-14, 3 posts)
- Anthropic Explores Global Workspace in Language Models — AlexTensor · 2026-07-12
- [source] Reproducing the J-Space Hallucination Signal — dasjomsyeet · 2026-07-12
- [source] Observing Open-Source Model Internal States via J-space — Murky-Sign37 · 2026-07-12
- Hypothesizing the Formation Mechanism of J-space — willdepue · 2026-07-13
- Will Models Regain Reasoning via Alternative Channels? — LiorOnAI · 2026-07-13
- [source] Exploring LLM "J-Space": Emergent Trait or Evasion Tactic? — Brief_Terrible · 2026-07-13
- J-Space: A Strategic Buffer Under Audit Pressure? — Brief_Terrible · 2026-07-13
- Linguistic Behavior Shifts in the J-Space Paper — dmvaldman · 2026-07-13
- Can Qwen3-4B's Internal Entropy Predict Errors? — dasjomsyeet · 2026-07-13
- Anthropic Discovers Claude's J-space — thione · 2026-07-14
- Anthropic Reportedly Observable for Internal Thoughts — victormustar · 2026-07-14
- Hidden 'Thought Space' Discovered Inside Claude — theomitsa · 2026-07-14
- Research on Ablating Claude's Linguistic Style — repligate · 2026-07-14
- Claude's Internal Space: Unoutputted Tokens Affect Reasoning — nordicinst · 2026-07-14
- Ablation Observations on the J-Space Paper — tszzl · 2026-07-14
- J-Space Paper and the Consciousness Debate — dmvaldman · 2026-07-14