Stanford's Matryoshka Attribution tops mech interp benchmark, removes refusal by resetting 1% of weights
stanfordnlp · x · 2026-09-26
Stanford researchers (Aryaman Arora, Noah Goodman, Dan Jurafsky, Christopher Potts, et al.) release Matryoshka Attribution (MAttr), a new interpretability method.
- Reframes attribution as finding nested subsets of internal components that minimize a downstream loss, using a differentiable sigmoid top-k mask operator trained across all sparsities simultaneously
- Ranks #1 on the official Mechanistic Interpretability Benchmark leaderboard, identifying sparse, task-transferable circuits
- Practical demo: MAttr trained with RL on refusal judge scores shows restoring just 1% of Llama 3.1 8B Instruct's weights to base state suffices to remove refusal behavior
Related event: Stanford's Matryoshka Attribution Tops Interpretability Benchmarks(2 posts)→
More from Research
- Claude pushes theoretical physics to nine loops, beating the previous eight-loop record — ctjlewis · 2026-09-26
- Reverse-engineering swarms: recover code from masked outputs for interpretable models — cephaloform · 2026-09-26
- One amphibian question can tell if a model treats your prompt as a capability eval — AdtRaghunathan · 2026-09-26
- Falcon Perception-HD GRPO update detects 100-500 objects, dropping NMS entirely — HildeKuehne · 2026-09-26
- An AI forecaster predicts misalignment from training data before training begins — tomekkorbak · 2026-09-26
- Mathematical archaeology: tracing where the AI proof that ζ(5) is irrational came from — ctjlewis · 2026-09-26