Circuit interpretability hits a wall as ablations miss across tasks
A COLM 2026 workshop paper finds extracted circuits lack specificity, with ablations harming other tasks; researchers explain why component-level specificity is low while neuron-level is higher, citing related SAE work on concept manifolds.
2026-09-05 ~ 2026-09-05 · 4 related posts
- Circuits Fail the Specificity Test: Ablating One Task's Circuit Hurts Others Equally, COLM Oral Finds — bearseascape · 2026-09-05
- Interpretability discussion: why component-level circuits lack specificity but neuron-level fares better — aryaman2020 · 2026-09-05
- Why MLP neurons split by output class: fixed residual-stream write directions explain it — aryaman2020 · 2026-09-05
- New Paper Asks Whether SAEs Capture Concept Manifolds, Finds a 'Dilution' Failure Mode — aryaman2020 · 2026-09-05