Attribution as a training objective: MAttr answers the Impossibility Theorems paper
aryaman2020 · x · 2026-09-24
The author frames MAttr's methodological stance: rather than axiomatically defined attribution outcomes, classic interpretability problems (like attribution) should be treated as training objectives optimized via gradient descent against the task of interest—directly answering the call of the "Impossibility Theorems" paper by Bilodeau, Jaques, Koh and Kim. The authors also see it as a more mechanistic complement to metamodels, sharing the same worldview.
More from Research
- MEMOIR benchmark: 117 synthetic oncology patients to test AI clinical memory — jefrankle · 2026-09-24
- 1.7h of intervention data beats 21h of demos: finetuning π0.5 on manufacturing — DominiqueCAPaul · 2026-09-24
- Fine-tuning π0.5 on a real factory task: data diversity beats raw scale, hitting 98% success — DominiqueCAPaul · 2026-09-24
- π0.5 finetuning on real manufacturing: 1 hour of clean data beat the previous 17 — DominiqueCAPaul · 2026-09-24
- Inside Anthropic's new molecular biology lab: Claude proposes, scientists verify — AnthropicAI · 2026-09-24
- Claude discovers unknown enzyme system in phage DNA with CRISPR-like repeats — AnthropicAI · 2026-09-24