Claude's Internal Space: Unoutputted Tokens Affect Reasoning

nordicinst · x · 2026-07-14

This repost highlights an article on Anthropic/Claude interpretability research. By peeking into Claude's internal space (J-space), researchers found that even tokens not directly outputted continuously influence the model's reasoning process.

The core insight is that the model doesn't think "mysteriously"; rather, its internal states and decision trajectories can be observed mathematically. These findings offer valuable reference for understanding model reasoning, interpretability, and alignment.

Related event: Anthropic Reveals Claude's Hidden Reasoning Space: J-space(16 posts)→

Original post →

More from Research

Research channel →