MAttr turns attribution into a training objective, hitting MIB SOTA in just 500 steps

aryaman2020 · x · 2026-09-24

Aryaman introduces MAttr, a new interpretability method that frames attribution itself as a training objective:

Related event: Stanford's Matryoshka Attribution Tops Mechanistic Interpretability Benchmarks at 2.9x Runner-up(7 posts)→

Original post →

More from Research

Research channel →