Interpretability's two halves: how the sausage is made vs what the sausage is
thebasepoint · x · 2026-09-27
thebasepoint sketches a rough split in interpretability research: "how the sausage is made" vs "what the sausage is". Empirically the latter has mattered far more for understanding model behavioral issues, while the former is more beautiful; he estimates NLA is 90% the latter, noting it won't reveal how addition splits into mod-10 and magnitude components.
Related event: AI safety researchers debate interpretability paths and the value of NLA(8 posts)→
More from Research
- Quail: open-source AI-SQL engine hits 1B+ tokens/min on a single H100 — sh_reya · 2026-09-27
- NYU's Tal Linzen cites two papers arguing tool use breaks Bender & Koller's 'no meaning' case — tallinzen · 2026-09-27
- DeepMind researcher argues AI-debate authors ignore empirical evidence that contradicts them — AndrewLampinen · 2026-09-27
- Google researcher Lampinen pens long thread rebutting the stochastic parrots argument on LLM meaning — AndrewLampinen · 2026-09-27
- Xiaomi publishes MiMo-V2.6 paper on scaling reinforcement learning toward LLM-Core — KyeGomezB · 2026-09-27
- The brain is a predictive machine: remove reality's correction signal and it hallucinates — alfcnz · 2026-09-27