Interp vs mech interp: you can explain models without hunting for mechanisms
voooooogel · x · 2026-09-06
A debate within the interpretability community: @voooooogel argues for calling the broader group just "interp" rather than "mech interp," since these approaches explain models without necessarily seeking internal mechanisms. Pushing back on the framing of mech interp as "curiosity driven" vs probes as purely pragmatic, the author contends that non-mechanistic interp can be just as curiosity driven—top-down rather than bottom-up—and that pragmatists reach for non-mech tools mainly because they currently work better.
More from Research
- AI editing a math paper spots a counterexample to a proposition the author planned to cite — jessi_cata · 2026-09-06
- Limits to narrow LLM complementarity: why 'taste is the bottleneck' won't hold — zetalyrae · 2026-09-06
- NeurIPS 2026 Registration Opens: Sydney Main Site Plus Atlanta and Paris Satellites — NeurIPSConf · 2026-09-06
- PolyU-Led Survey Maps Human-Centric AI: 3 Perspectives, 6 Layers from Body Perception to Embodied Agents — 机器之心 · 2026-09-06
- Terence Tao Clarifies Navier-Stokes Rumor Was Misread Hypothetical; Clay Still Lists Problem Unsolved — 机器之心 · 2026-09-06
- Paper: One training query recovers 72% of full-data on-policy distillation gains — heghbalz · 2026-09-06