Matryoshka Attribution: reverting 1% of weights removes most refusal in Llama 3.1 8B

aryaman2020 · x · 2026-09-29

The paper "Matryoshka Attribution" introduces a method that traces model behaviors back to the smallest set of internal components responsible for them.

This suggests safety alignment may be concentrated in a tiny fraction of weights, with direct implications for interpretability and safety research.

Original post →

More from Research

Research channel →