Attribution as a training objective: MAttr answers the Impossibility Theorems paper

aryaman2020 · x · 2026-09-24

The author frames MAttr's methodological stance: rather than axiomatically defined attribution outcomes, classic interpretability problems (like attribution) should be treated as training objectives optimized via gradient descent against the task of interest—directly answering the call of the "Impossibility Theorems" paper by Bilodeau, Jaques, Koh and Kim. The authors also see it as a more mechanistic complement to metamodels, sharing the same worldview.

Related event: Stanford's Matryoshka Attribution Tops Mechanistic Interpretability Benchmarks at 2.9x Runner-up(7 posts)→

Original post →

More from Research

Research channel →