Matryoshka Attribution finds neural network circuits via gradient descent, tops interpretability benchmark by 2.9x
stanfordnlp · x · 2026-09-26
Noah Goodman highlighted a new paper introducing Matryoshka Attribution (MAttr), an attribution method that uses gradient descent to locate which parts of a neural network drive a given behavior.
MAttr ranks #1 on the Mechanistic Interpretability Benchmark by a wide margin, scoring 2.9× the runner-up.
Related event: Stanford's Matryoshka Attribution Tops Interpretability Benchmarks(2 posts)→
More from Research
- An AI forecaster predicts misalignment from training data before training begins — tomekkorbak · 2026-09-26
- Mathematical archaeology: tracing where the AI proof that ζ(5) is irrational came from — ctjlewis · 2026-09-26
- 2-Layer Recurrent Networks Match 32-Layer Feedforward Baselines at Same Compute, Thread Claims — mike64_t · 2026-09-26
- Richard Socher's new book 'The Eureka Machine' argues AI unlocks a new era of science — RichardSocher · 2026-09-26
- AI-drafted 166-page Navier-Stokes proof is correct but nearly unreadable for humans — Pascallisch · 2026-09-26
- QuackIR: Jimmy Lin's EMNLP Paper Shows RDBMSes Match Vector DBs for RAG Retrieval — lintool · 2026-09-26