MAttr traces Llama 3.1 refusals to 1% of parameter delta, resetting them disables refusal

aryaman2020 · x · 2026-09-24

MAttr works beyond representations: it can trace qualitative LLM behaviors to a subset of parameters updated during training, by (1) training attribution scores with RL using a judge score (StrongREJECT) as reward, and (2) applying the mask to model weights across two checkpoints instead of representations across prompts. Result: Llama 3.1 8B refusals map to 1% of the parameter delta between Base and Instruct; resetting those weights in the Instruct model keeps its capabilities while largely turning off refusal.

Related event: Stanford Team's Matryoshka Attribution Tops Mechanistic Interpretability Benchmarks(6 posts)→

Original post →

More from Research

Research channel →