Yoav Goldberg: interpretability that serves steering is just steering research with a handicap
yoavgo · x · 2026-08-26
Yoav Goldberg argues sharply: if the goal of interpretability research is steering, and interpretability methods are judged by how well they improve steering, then it is no longer interpretability research — it is steering research with a useless handicap (using interp methods). He can agree that "interpretability should be causal," but says "actionable" implies we care about the effect itself, not the explanation.
More from Research
- LangChain Introduces WikiBench to Evaluate Documentation Agents — LangChain · 2026-08-26
- Rotations form curved spaces; exponentials resolve singularities — FrnkNlsn · 2026-08-26
- GeneralistAI demonstrates GEN-1.5 physical prompt steerability — E0M · 2026-08-26
- How AI Agents Work Part 1: Guessing the Next Token — sethjuarez · 2026-08-26
- Interpreting AI Models via Topology and Differential Algebra on N-Dimensional Spheres — _xjdr · 2026-08-26
- Looking back at 2017's 'Generating Sentences by Editing Prototypes' — srush_nlp · 2026-08-26