Can Internal J-Space in Models Be Reverse-Trained?
Intelligent_Gear5739 · reddit · 2026-07-11
The post references an Anthropic video and paper detailing a J-Space within the model that acts like "cached thinking." Certain words here are tied to the model’s reasoning process, and removing them degrades performance.
The author poses two further questions: Could we train models backward from a target J-Space state to force convergence toward specific internal states? Furthermore, does this mean we could suppress the model from "thinking" about certain concepts or make others appear more frequently in the J-Space?
Related event: Anthropic's J-Space Research Illuminates Claude's Inner Reasoning(3 posts)→
More from Research
- New paper defines self-state attacks, showing OS defenses leave four agent-memory cases indistinguishable — Justgototheeffinmoon · 2026-07-22
- Krea 2 users recommend a two-pass Clownshark sampler setup for sharper image details — listopalafoto · 2026-07-22
- Animation shows how an MLP’s first-layer weights change while learning MNIST — CatAstro_Piyush · 2026-07-22
- Project APE finds verifier reliability drops when papers contain multiple errors — soumitrashukla9 · 2026-07-22
- Project APE says verifier costs fell about 90x in a year as Chinese open models lead — soumitrashukla9 · 2026-07-22
- OpenAI-linked paper says capability RL can make models more reward-seeking — MariusHobbhahn · 2026-07-22