Will interpretability ever be "solved"? A researcher argues probably not

burny_tech · x · 2026-09-05

The author argues interpretability will likely never be "solved": full interp would mean predicting any behavior, activation pattern, and intervention outcome of any model — any architecture, learning algorithm, or dataset — under the smallest-Kolmogorov theory, which seems impossible. Still, echoing physics and biology, no science is ever fully solved; better interp science can asymptotically approach that ideal. Predictive and explanatory power remain the core yardsticks.

Original post →

More from AGI Musings

AGI Musings channel →