Yoav Goldberg Sparks Debate: Mechanistic Understanding Should Trump Steering in Interpretability Research
On August 26, prominent NLP scholar Yoav Goldberg fired off a series of posts on the positioning of interpretability research, clashing with multiple users and igniting a debate over "what interpretability is actually for." The core divide: should interpretability research aim to understand model mechanisms, or be judged by downstream-task performance (steering, probing, intervention)?
Confirmed
- Goldberg argues the core goal of interpretability research is further understanding of model mechanisms; inventing new methods is a side effect. If the goal is mechanistic understanding, then progress that aids understanding is real.
- He contends that if the research goal is "steering" and it's judged by steering effectiveness, it is essentially no longer interpretability research but steering research burdened with useless shackles (interpretability methods); similarly, he calls "actionable interpretability research" an inferior form of interpretability, and provocatively invited pushback.
- He adds that some explanations may lack causal power yet remain valuable — e.g., research showing knowledge representations in model weights are highly redundant; while causal experiments are hard to design, the conclusions still matter. Using "knowledge editing" as an example, he cautions that concluding knowledge is stored at a location merely because editing weights there changes behavior is misleading, since the intervention may hit the last link in a pathway rather than where knowledge actually resides.
- @EEnouen countered: interpretability research should hill-climb toward specific downstream goals (steering, probing) rather than optimize intrinsic metrics like faithfulness; if an explanation method can't beat simple baselines at predicting behavior, intervening, or steering, it may be "fake progress."
- @aryaman2020 agreed and added that the main flaw of "practical interpretability" is that it tackles the same problems as post-training researchers, who aren't constrained to interpretability as a specific method.
Why it matters
The dispute cuts to the field's current路线之争: as steering/manipulation becomes a hot application direction, whether engineering-oriented evaluation criteria are eroding deep mechanistic understanding or instead providing a reality check against "fake progress" will shape researchers' method choices and how evaluation systems are built.
2026-08-26 ~ 2026-08-26 · 8 related posts
Primary sources
- [source] Yoav Goldberg: actionable interpretability research is the lesser form — yoavgo · 2026-08-26
- "Actionable interpretability" argued to be an inferior form of research — aryaman2020 · 2026-08-26
- Interpretability is "false progress" if it cannot improve intervention capabilities — EEnouen · 2026-08-26
- [source] Interpretability methods should hill-climb on downstream goals, not intrinsic metrics — EEnouen · 2026-08-26
- Yoav Goldberg: interpretability that serves steering is just steering research with a handicap — yoavgo · 2026-08-26
- [source] Yoav Goldberg: Interpretability Core is Understanding, Not Just Methods — yoavgo · 2026-08-26
- Yoav Goldberg: Interpretability core is understanding mechanisms, not just methods — yoavgo · 2026-08-26
- Knowledge editing experiments may mislead interpretability conclusions — yoavgo · 2026-08-26