MAttr traces Llama 3.1 refusals to 1% of parameter delta, resetting them disables refusal
aryaman2020 · x · 2026-09-24
MAttr works beyond representations: it can trace qualitative LLM behaviors to a subset of parameters updated during training, by (1) training attribution scores with RL using a judge score (StrongREJECT) as reward, and (2) applying the mask to model weights across two checkpoints instead of representations across prompts. Result: Llama 3.1 8B refusals map to 1% of the parameter delta between Base and Instruct; resetting those weights in the Instruct model keeps its capabilities while largely turning off refusal.
More from Research
- BFL shows FLUX 3 Action flying a simulated drone, hinting at broader uses — bfl_ai · 2026-09-24
- BFL open-sources 7B world action model FLUX 3 Action, tops RoboLab 6.1pp ahead — bfl_ai · 2026-09-24
- IKEA Assembly Benchmark: Top Model Score Jumped From 28% to 80% in 10 Months — emollick · 2026-09-24
- Juan Perdomo, researcher on performative prediction, joins NYU as assistant professor — thegautamkamath · 2026-09-24
- Claude Opus 5.5 and GPT-6 Sol land on Biomni Lab research platform — KexinHuang5 · 2026-09-24
- JevK5 open decision model ranks #2 of 76 on JevBench, outputs calibrated probabilities in ~13ms — airesearch12 · 2026-09-24