Matryoshka Attribution: nested gradient-descent attribution tops interpretability benchmark by 2.9x

ChrisGPotts · x · 2026-09-24

A new paper introduces Matryoshka Attribution (MAttr), an attribution method that uses gradient descent to find which parts of a neural network drive a behavior, learning a nested ranking of model components across sparsities and directly optimizing intervention performance.

MAttr ranks #1 on the Mechanistic Interpretability Benchmark by a wide margin — 2.9× the runner-up — transfers across tasks, and extends to parameter attribution. Co-authors include Chris Potts, Dan Jurafsky, and Noah Goodman.

Related event: Stanford's Matryoshka Attribution Tops Mechanistic Interpretability Benchmarks at 2.9x Runner-up(7 posts)→

Original post →

More from Research

Research channel →