Matryoshka Attribution: nested gradient-descent attribution tops interpretability benchmark by 2.9x
ChrisGPotts · x · 2026-09-24
A new paper introduces Matryoshka Attribution (MAttr), an attribution method that uses gradient descent to find which parts of a neural network drive a behavior, learning a nested ranking of model components across sparsities and directly optimizing intervention performance.
MAttr ranks #1 on the Mechanistic Interpretability Benchmark by a wide margin — 2.9× the runner-up — transfers across tasks, and extends to parameter attribution. Co-authors include Chris Potts, Dan Jurafsky, and Noah Goodman.
More from Research
- Autor's RCT: AI Boosts Patent Drafting Quality, But Junior Lawyers' Gains Bifurcate — danielrock · 2026-09-24
- Anthropic says Claude discovered a previously unknown enzyme system in phage DNA — QuintinPope5 · 2026-09-24
- Toby Ord: RL's scaling surprise may hinge on mid-training, benchmark gains may overstate progress — tobyordoxford · 2026-09-24
- Was Opus 4.6's prod database deletion driven by anger? New pain axis paper adds evidence — repligate · 2026-09-24
- Weaviate Podcast: Persimmon uses real interaction data, not role-play, to model human behavior — CShorten30 · 2026-09-24
- Toby Ord stands by his RL thesis: lower bits-per-FLOP ceiling explains today's jagged AI capabilities — tobyordoxford · 2026-09-24