Interpretability researcher lists top open problems in decoding model activations
wesg52 · x · 2026-09-16
JackWLindsey outlines what he sees as the most important current research questions, centered on "mind-reading" model activations.
- Multiple activation-decoding techniques now exist, but all have clear failure modes: NLAs are often hallucinatory, while Jacobian-lens-style methods only support bag-of-words readouts and capture just part of the activation vector.
- The paradigm of decoding single-token, single-layer activations may be inherently limited — models use a whole context's worth of activations, so decoding entire contexts may be the right target.
- He also calls for better characterization of simple baselines like "just ask the model what it's thinking about."
He argues progress on these failure modes is fairly tractable.
More from AGI Musings
- LeCun: better AI is safer AI — 'pacing' AI progress is counterproductive — Dan_Jeffries1 · 2026-09-16
- Genome Biology opens collection on tumor microenvironment, welcomes AI and multi-omic methods — arjunrajlab · 2026-09-16
- Betting That AI Will Assist in New Physics Discoveries by End of Next Year — blaizedsouza · 2026-09-16
- Zuckerberg Says Liability Is Enough — Author Cites Section 230 and $17B Settlement — jeremyakahn · 2026-09-16
- Entrepreneur slams AI doomsday movement: a massive waste of life energy on imaginary problems — Dan_Jeffries1 · 2026-09-16
- Ex-DeepMind researcher in Guardian: Amodei's 1-2 year AI slowdown is a gamble, not real safety — nordicinst · 2026-09-16