Anthropic Releases Full Paper on LLM "J-Space"

wesg52 · x · 2026-07-07

Anthropic's interpretability team has officially released the full research paper on the language model "Judgment Space" (J-space). J-space is a small subset of model activations that reflects the model's internal state, emotions, implicit thoughts, and safety awareness, existing independently of output tokens. The research reveals a systematic distinction between a model's "reportable cognition" and "implicit internal processing," opening a new direction for AI interpretability and safety research. The paper was co-authored by Jack Lindsey, Nick Sofroniew, and the Anthropic interpretability team.

Related event: Anthropic Discovers Global Workspace Inside Claude(102 posts)→

Original post →

More from Research

Research channel →