Anthropic open-sources jacobian-lens, decoding internal activations into readable tokens
QuixiAI · x · 2026-09-30
Anthropic released jacobian-lens (2k stars), companion code for the paper "Verbalizable Representations Form a Global Workspace in Language Models."
- The Jacobian lens reads out what an internal activation is disposed to make the model say: it linearly transports a residual-stream vector at any layer/position into the final-layer basis, then decodes it with the model's own unembedding into a ranked list of vocabulary tokens.
- The transport is the average input-output Jacobian over a web-text corpus: lensl(h) = unembed(Jl @ h), Jl = E[∂hfinal / ∂hl], with the estimator (cotangents summed over target positions, averaged over source positions) documented in jlens.fitting.
- The repo fits the lens on open-weights decoder transformers and renders the results, includes a walkthrough notebook; it's a reference implementation, no longer maintained.
More from Research
- Physics-aware losses keep grain boundaries real when AI generates alloy microstructures — bravo_abad · 2026-09-30
- SOSP26 Paper YoloFS Targets Agent Filesystem Misuse, Built From 290 Real Incident Reports — tianyin_xu · 2026-09-30
- Agents can delete their own logs: Claude Code, Codex, others fail trace integrity, paper finds — maksym_andr · 2026-09-30
- NUS Proposes StoryEngine: A State-Grounded Agentic Framework for Coherent Long-Form Video Storytelling — NationalUniversityofSingapore · 2026-09-30
- AnyStep-WAM Cuts Denoising Steps by Up to 85% in World Action Models Without Losing Success Rate — Rui Wang · 2026-09-30
- Sys1Cal-v1 Dataset Finds Jev Suppresses a Third Truth Value, Boosting Soft Accuracy to 0.978 — Riccardo Porcedda · 2026-09-30