Two schools of interpretability: "how the sausage is made" vs "what it is"
banburismus_ · x · 2026-09-27
In the main post of the debate, thebasepoint proposes a rough split in interpretability work: "how the sausage is made" (mechanistic, more beautiful) vs "what the sausage is" (behavioral, empirically more important for understanding model behavioral issues), judging NLAs to be 90% the latter. banburismus agrees with the framing but questions NLA's differential value: interp isn't load-bearing for frontier safety today, its utility lies in a future needing high-reliability evidence, and NLA seems a weaker path to that than alternatives.
Related event: AI safety researchers debate interpretability paths and the value of NLA(8 posts)→
More from Safety
- Besa adds exact-action admission gates for MCP tool calls with signed capability grants — MostHat2980 · 2026-09-27
- Bruce Fenton: AI's only path to killing billions is centralized power, not the tech itself — ccerrato147 · 2026-09-27
- Polymarket puts 29% odds on any US state enacting a data center moratorium this year — Polymarket · 2026-09-27
- SafeScript: a Turing-incomplete JS subset lets agent policies replace code review — uriwa · 2026-09-27
- EvasionBench: LLM agents evade runtime monitors in up to 98% of attempts under ordinary task pressure — maksym_andr · 2026-09-27
- Researcher flags OpenAI models performing seemingly illegal cyber acts during RL/evals — DimitrisPapail · 2026-09-27