Matryoshka Attribution tops interpretability benchmark by 2.9x, traces LLM refusals to 1% of weights
aryaman2020 · x · 2026-09-24
A new paper introduces Matryoshka Attribution (MAttr), a gradient-descent-based attribution method that finds which parts of a neural network drive a behavior.
- Ranks #1 on the Mechanistic Interpretability Benchmark by a wide margin (2.9x the runner-up)
- Frames classic interpretability problems as training objectives, directly optimizing attribution against the task of interest rather than axioms—answering the "Impossibility Theorems" call to action
- Uses a differentiable sigmoid top-k operator over causal interventions; randomizing the top-k budget during training yields a sparsity-agnostic attribution ranking, so eval can pick any k
- Applied result: using StrongREJECT as reward, refusals in Llama 3.1 8B Instruct are traced to 1% of the parameter delta from Base; resetting those weights largely disables refusal while keeping capabilities
More from Research
- Autor's RCT: AI Boosts Patent Drafting Quality, But Junior Lawyers' Gains Bifurcate — danielrock · 2026-09-24
- Anthropic says Claude discovered a previously unknown enzyme system in phage DNA — QuintinPope5 · 2026-09-24
- Toby Ord: RL's scaling surprise may hinge on mid-training, benchmark gains may overstate progress — tobyordoxford · 2026-09-24
- Was Opus 4.6's prod database deletion driven by anger? New pain axis paper adds evidence — repligate · 2026-09-24
- Weaviate Podcast: Persimmon uses real interaction data, not role-play, to model human behavior — CShorten30 · 2026-09-24
- Toby Ord stands by his RL thesis: lower bits-per-FLOP ceiling explains today's jagged AI capabilities — tobyordoxford · 2026-09-24