MAttr technical details: differentiable sigmoid top-k merges causal interventions and mask learning

aryaman2020 · x · 2026-09-24

Technical notes on MAttr: it combines the strengths of causal interventions, gradient-based attributions and mask learning, using a simple differentiable sigmoid top-k operator (no sparsity loss, no straight-through tricks) to parametrize causal interventions. Randomizing the top-k budget during training induces an attribution ranking; at eval time you simply set k to any desired sparsity.

Related event: Stanford's Matryoshka Attribution Tops Mechanistic Interpretability Benchmarks at 2.9x Runner-up(7 posts)→

Original post →

More from Research

Research channel →