Matryoshka Attribution finds neural network circuits via gradient descent, tops interpretability benchmark by 2.9x

stanfordnlp · x · 2026-09-26

Noah Goodman highlighted a new paper introducing Matryoshka Attribution (MAttr), an attribution method that uses gradient descent to locate which parts of a neural network drive a given behavior.

MAttr ranks #1 on the Mechanistic Interpretability Benchmark by a wide margin, scoring 2.9× the runner-up.

Related event: Stanford's Matryoshka Attribution Tops Interpretability Benchmarks(2 posts)→

Original post →

More from Research

Research channel →