Stanford's Matryoshka Attribution tops mech interp benchmark, removes refusal by resetting 1% of weights

stanfordnlp · x · 2026-09-26

Stanford researchers (Aryaman Arora, Noah Goodman, Dan Jurafsky, Christopher Potts, et al.) release Matryoshka Attribution (MAttr), a new interpretability method.

Related event: Stanford's Matryoshka Attribution Tops Interpretability Benchmarks(2 posts)→

Original post →

More from Research

Research channel →