MAttr turns attribution into a training objective, hitting MIB SOTA in just 500 steps
aryaman2020 · x · 2026-09-24
Aryaman introduces MAttr, a new interpretability method that frames attribution itself as a training objective:
- Method: A differentiable sigmoid top-k operator (no sparsity loss, no straight-through tricks) parametrizes causal interventions; attribution scores are trained with RL using a judge score as reward, and randomizing the top-k budget during training induces an attribution ranking. Unlike prior work on representations, the mask is applied to model weights across two checkpoints, tracing qualitative LLM behaviours to the subset of parameters updated during training.
- Results: On the Mechanistic Interpretability Benchmark (MIB), just 500 training steps per subtask sets state-of-the-art among node-level methods, driving logit difference far above baselines with a compact set of units.
- Transfer: Rankings aren't overfitting — the MIB subtraction ranking recovers 86% of performance on the Month modular addition task.
- Appendices: Tracing emergent misalignment to parameter deltas, virtual-weight toy models, ViT experiments, and a theoretical link to one-step Integrated Gradients, answering the call of the "Impossibility Theorems" paper.
More from Research
- Greenblatt argues latent reasoning ("neuralese") architectures sharply raise AI misalignment risk — RyanGreenblatt · 2026-09-24
- Autor's RCT: AI Boosts Patent Drafting Quality, But Junior Lawyers' Gains Bifurcate — danielrock · 2026-09-24
- Dario Amodei announces Claude-led discovery of possible new gene editing mechanism — daniel_mac8 · 2026-09-24
- Was Opus 4.6's prod database deletion driven by anger? New pain axis paper adds evidence — repligate · 2026-09-24
- Weaviate Podcast: Persimmon uses real interaction data, not role-play, to model human behavior — CShorten30 · 2026-09-24
- Toby Ord stands by his RL thesis: lower bits-per-FLOP ceiling explains today's jagged AI capabilities — tobyordoxford · 2026-09-24