Two schools of interpretability: "how the sausage is made" vs "what it is"

banburismus_ · x · 2026-09-27

In the main post of the debate, thebasepoint proposes a rough split in interpretability work: "how the sausage is made" (mechanistic, more beautiful) vs "what the sausage is" (behavioral, empirically more important for understanding model behavioral issues), judging NLAs to be 90% the latter. banburismus agrees with the framing but questions NLA's differential value: interp isn't load-bearing for frontier safety today, its utility lies in a future needing high-reliability evidence, and NLA seems a weaker path to that than alternatives.

Related event: AI safety researchers debate interpretability paths and the value of NLA(8 posts)→

Original post →

More from Safety

Safety channel →