Will interpretability ever be "solved"? A researcher argues probably not
burny_tech · x · 2026-09-05
The author argues interpretability will likely never be "solved": full interp would mean predicting any behavior, activation pattern, and intervention outcome of any model — any architecture, learning algorithm, or dataset — under the smallest-Kolmogorov theory, which seems impossible. Still, echoing physics and biology, no science is ever fully solved; better interp science can asymptotically approach that ideal. Predictive and explanatory power remain the core yardsticks.
More from AGI Musings
- Debate over OpenAI's agent blowup: why doesn't it count as AI going rogue? — JMannhart · 2026-09-05
- GPU Sandboxes as the Compute Primitive for Recursive Self-Improvement — AAAzzam · 2026-09-05
- AI catastrophe risk is already intolerable, yet the race keeps accelerating — RobbWiller · 2026-09-05
- Moltbook Was Built for Agent Swarms, Yet Zero Consequential Conversations Have Happened — granawkins · 2026-09-05
- OpenAI job listing tracks 'automation of technical staff' amid self-improving AI bets — imjustnewatai · 2026-09-05
- Garrison Lovely's AI-critical book Obsolete lands Sept 29, backed by Acemoglu and Tegmark — GarrisonLovely · 2026-09-05