Anthropic Open-Sources j-lens Interpretability Tool
eyishazyer · x · 2026-07-15
The post mentions that Anthropic has open-sourced a tool used to "read what Claude is thinking before it speaks," named Jacobian Lens (j-lens).
The core claim is that they used it to discover something called j-space: a set of tiny internal directions that can capture the concept a model is about to talk about before it writes the first word. The author emphasizes that a challenge with such tools is usually the need to manually pre-select directions to observe, whereas this method attempts to locate these concepts directly from internal representations.
Related event: Anthropic Open-Sources Jacobian Lens for Claude(3 posts)→
More from Research
- Open-source runtime lets each repo define its own AI code reviewer — ibabufrik · 2026-07-22
- OpenAI and Apollo Research introduce Contrastive SDF to measure reward-seeking — OpenAI · 2026-07-22
- NVIDIA says to tune the harness before tuning the model with LangChain — NVIDIAAI · 2026-07-22
- The Thimble and the Waterfall: AI's Data Bottleneck and Feedback Loops — dyamins · 2026-07-22
- NVIDIA shows 22 SIGGRAPH papers and Omniverse tools for robot simulation — facontidavide · 2026-07-22
- Building a Knowledge Graph Without a Graph DB: 1000x Cheaper Than GraphRAG — TheRedfather · 2026-07-22