Anthropic Releases Full Paper on LLM "J-Space"
wesg52 · x · 2026-07-07
Anthropic's interpretability team has officially released the full research paper on the language model "Judgment Space" (J-space). J-space is a small subset of model activations that reflects the model's internal state, emotions, implicit thoughts, and safety awareness, existing independently of output tokens. The research reveals a systematic distinction between a model's "reportable cognition" and "implicit internal processing," opening a new direction for AI interpretability and safety research. The paper was co-authored by Jack Lindsey, Nick Sofroniew, and the Anthropic interpretability team.
Related event: Anthropic Discovers Global Workspace Inside Claude(102 posts)→
More from Research
- Nature paper images cellular activity across all organs, revealing body-wide circuits — arjunrajlab · 2026-09-11
- SignNet 1M Dataset Released for Sign Language Research — ducha_aiki · 2026-09-11
- ECCV26 Oral: Flow Matching Enables Single-Stage Multi-View Point Cloud Registration — ducha_aiki · 2026-09-11
- InFlux++ Method Released — ducha_aiki · 2026-09-11
- Skyfall GS Uses Flux to Refine Gaussian Splatting, Accepted at ECCV 2026 — ducha_aiki · 2026-09-11
- Could 10k agents discover learning methods beyond backprop, or just tweak existing ones? — SeunghyunSEO7 · 2026-09-11