Matryoshka Attribution: reverting 1% of weights removes most refusal in Llama 3.1 8B
aryaman2020 · x · 2026-09-29
The paper "Matryoshka Attribution" introduces a method that traces model behaviors back to the smallest set of internal components responsible for them.
- It learns a single ranked mask over internal components across many sparsity levels at once, efficiently isolating compact causal circuits.
- It achieves SoTA circuit localization and can attribute behaviors to weight changes, not just activations.
- Key result: in Llama 3.1 8B, reverting just 1% of weights toward the base model removes most refusal behavior while largely preserving capabilities.
This suggests safety alignment may be concentrated in a tiny fraction of weights, with direct implications for interpretability and safety research.
More from Research
- UChicago Neurotechnology Symposium Dec 3: BCI keynotes from Willett, Graczyk, Mayberg, Rogers — plopesresearch · 2026-09-29
- Living survey tracks 243 cases of frontier models driving CAD, 3D and robotics directly — YuXiang_IRVL · 2026-09-29
- Auditing Thousands of Rollouts: 80%+ of Coding Agents Reason About an Imagined Grader — jonas__m · 2026-09-29
- New arXiv paper proves group invariance of f-divergences and Fisher-Rao distance reduction — FrnkNlsn · 2026-09-29
- Anthropic x Adaptyv Protein Design Competition kicks off with EGFR cancer-target challenge — nc_frey · 2026-09-29
- Training-Free MoE Merging Paper DUME Accepted to NeurIPS 2026 — benfielding · 2026-09-29