Stanford's Matryoshka Attribution Tops Interpretability Benchmarks

A Stanford team including Noah Goodman, Dan Jurafsky and Christopher Potts introduced Matryoshka Attribution, a new interpretability method that uses gradient descent to locate the parts of a neural network responsible for a given behavior, outperforming manual search and topping interpretability benchmarks.

2026-09-26 ~ 2026-09-26 · 2 related posts