Can Internal J-Space in Models Be Reverse-Trained?

Intelligent_Gear5739 · reddit · 2026-07-11

The post references an Anthropic video and paper detailing a J-Space within the model that acts like "cached thinking." Certain words here are tied to the model’s reasoning process, and removing them degrades performance.

The author poses two further questions: Could we train models backward from a target J-Space state to force convergence toward specific internal states? Furthermore, does this mean we could suppress the model from "thinking" about certain concepts or make others appear more frequently in the J-Space?

Related event: Anthropic's J-Space Research Illuminates Claude's Inner Reasoning(3 posts)→

Original post →

More from Research

Research channel →