Matryoshka Attribution tops interpretability benchmark by 2.9x, traces LLM refusals to 1% of weights

aryaman2020 · x · 2026-09-24

A new paper introduces Matryoshka Attribution (MAttr), a gradient-descent-based attribution method that finds which parts of a neural network drive a behavior.

Related event: Stanford's Matryoshka Attribution Tops Mechanistic Interpretability Benchmarks at 2.9x Runner-up(7 posts)→

Original post →

More from Research

Research channel →