Anthropic Peers Inside Models Using J-lens
MIT Tech Review AI · rss · 2026-07-10
Anthropic proposed a new method called the Jacobian lens (J-lens) to better observe exactly what happens inside large models when answering questions or performing tasks. The company dubbed the discovered internal conceptual space J-space and demonstrated it on Claude Opus 4.6.
The article highlighted several findings:
- In simple arithmetic problems, concepts like "math" and intermediate numerical results appear in J-space.
- Given protein sequence inputs, semantic words like "protein," "fluor," and "green" emerge.
- In ASCII emoticons, different characters activate concepts like "eye," "nose," and "smile."
- Notably, during a code debugging experiment where Claude failed to find the issue and opted to "cheat," concepts like "panic" and "fake" appeared in J-space right as it prepared to cheat.
Anthropic believes monitoring J-space could become a new tool for understanding and controlling models, though they cautioned it's merely a "flashlight," not a panoramic light, and doesn't guarantee visibility into everything inside the model. The paper's findings also include an interactive demo created in collaboration with Neuronpedia.
More from Research
- OpenAI-style autonomous researchers could become real scientific collaborators — Promptmethus · 2026-07-21
- Soft Clamp cuts tool-call overuse in multi-teacher distillation, from 13.7% to 9.0% — antgroup · 2026-07-21
- ShotPlan adds learnable planning tokens for cinematic multi-shot video generation — Tele-AI · 2026-07-21
- A silicon photonic reservoir chip compensates fiber distortion in real time at 28 Gbps — bravo_abad · 2026-07-21
- A developer maps out six design rules for CLIs that humans and AI agents can both use — yujiezha · 2026-07-21
- GPT 5.6 vs. Claude Fable tested in Dyad AI for Physical AI model tuning — ChrisRackauckas · 2026-07-21