Interpretability is going 'much worse than expected,' researcher says, citing missed 2026 milestone

stuhlmueller · x · 2026-09-03

Interpretability researcher stuhlmueller says the field is progressing much worse than he expected: in May 2023 he put 50/50 odds on extracting novel algorithmic insights from LLMs by end of 2026, but "we're nowhere close." Commenter aryaman pushes back, arguing induction heads, function vectors, and attention sinks already qualify.

Related event: Researchers Clash Over Whether Interpretability Is Delivering Real Insights(3 posts)→

Original post →

More from AGI Musings

AGI Musings channel →