Interpretability is going 'much worse than expected,' researcher says, citing missed 2026 milestone
stuhlmueller · x · 2026-09-03
Interpretability researcher stuhlmueller says the field is progressing much worse than he expected: in May 2023 he put 50/50 odds on extracting novel algorithmic insights from LLMs by end of 2026, but "we're nowhere close." Commenter aryaman pushes back, arguing induction heads, function vectors, and attention sinks already qualify.
Related event: Researchers Clash Over Whether Interpretability Is Delivering Real Insights(3 posts)→
More from AGI Musings
- A server-locked AI agent named Cairn changes the physical world through strangers' hands — No_Departure_9908 · 2026-09-03
- One month after predicting Singularity at 2032, Derya moves it up to 2030 — DeryaTR_ · 2026-09-03
- AI safety's 'meta work' questioned: what are hundreds of trainees actually doing? — austinc3301 · 2026-09-03
- Counterfactual: without reasoning models, AI today might just be reaching o3-level — Jsevillamol · 2026-09-03
- Why 'Hey Claude, watch Love Island for me' will never work: the case for personal agents — manosaie · 2026-09-03
- Central bankers are seriously discussing AI understanding monetary policy better than humans — VraserX · 2026-09-03