Stanford's Matryoshka Attribution tops mech interp benchmark at 2.9x the runner-up

aryaman2020 · x · 2026-09-24

A Stanford team (including Chris Potts and Dan Jurafsky) released Matryoshka Attribution (MAttr), a new attribution method that uses gradient descent over a sigmoid top-k mask to identify which internal variables of a neural network are responsible for a behavior.

Key points:

Related event: Stanford's Matryoshka Attribution Tops Mechanistic Interpretability Benchmarks at 2.9x Runner-up(7 posts)→

Original post →

More from Research

Research channel →