The J-lens Explained: Reading and Rewriting LLMs' Unspoken Concepts
CatAstro_Piyush · x · 2026-09-10
A blog post unpacks Anthropic's 'Verbalizable Representations Form a Global Workspace' paper: using a Jacobian lens, researchers find unspoken intermediate concepts in the residual stream — swapping a 'spider' direction for 'ant' flips answers from 8 legs to 6 in 54% of trials on Haiku 4.5 and 70% on Sonnet 4.5/Opus 4.5, with controls ruling out shortcut explanations. Covers the J-lens microscope, sparse J-space, and the read/write workflow.
More from Research
- Eric Drexler on the Hugging Face incident: system structure, not alignment, prevents AI collusion — sebkrier · 2026-09-10
- David Chalmers talks Anthropic's j-space and global workspace theory aboard a boat in the Galapagos — PeterBowdenLive · 2026-09-10
- A visual deep-dive catalogs 42+ representations of 3D, praised by HF engineer — pcuenq · 2026-09-10
- Ben Recht's forecasting lecture argues probability conflates frequency and belief — beenwrekt · 2026-09-10
- Spiced self-play accepted at CoRL: just 30 minutes of human data biases agents to right conventions — EugeneVinitsky · 2026-09-10
- Understanding FlashAttention: A Handbook Tracing FA1 to FA4 and Why HBM Traffic, Not FLOPs, Is the Bottleneck — techNmak · 2026-09-10