Interpretability's two halves: how the sausage is made vs what the sausage is

thebasepoint · x · 2026-09-27

thebasepoint sketches a rough split in interpretability research: "how the sausage is made" vs "what the sausage is". Empirically the latter has mattered far more for understanding model behavioral issues, while the former is more beautiful; he estimates NLA is 90% the latter, noting it won't reveal how addition splits into mod-10 and magnitude components.

Related event: AI safety researchers debate interpretability paths and the value of NLA(8 posts)→

Original post →

More from Research

Research channel →