Circuits Fail the Specificity Test: Ablating One Task's Circuit Hurts Others Equally, COLM Oral Finds
bearseascape · x · 2026-09-05
The author announces their interpretability paper has been accepted as an oral at the Sci-FM Workshop at COLM 2026, with new experiments since the original thread.
- Core finding: circuits extracted to explain model behavior fail a basic check — ablating one task's circuit hurts another task about as much as ablating that task's own circuit.
- Earlier results: component-level circuits were consistent and causally necessary, but not task-specific. The authors speculated superposition causes individual components to play multiple unrelated roles.
- New experiments: using better attribution methods (EAP-IG, RelP), the team dropped to the level of individual MLP neurons, motivated by recent work suggesting neurons already provide a sparse basis without extra training (avoiding SAE/CLT issues like feature splitting).
- New trade-off: neuron-level circuits are much more specific — ablating a task's own circuit hurts it far more — but far less consistent, with very few neurons recurring across examples of the same task.
The work raises fundamental doubts about whether circuit analysis truly captures how models solve tasks.
Related event: Circuit interpretability hits a wall as ablations miss across tasks(4 posts)→
More from Research
- IndianRailwayBench ranks LLMs by their ability to book tatkal train tickets — Paimaamu · 2026-09-06
- NEAR AI's open-source Lean agent solves all of Putnam Bench for just $111 — lukaszkaiser · 2026-09-06
- Russian startup Mostik bridges LLM hidden states, cutting cost to 1/20 — 机器之心 · 2026-09-06
- KV Cache Explained: Why It's Crucial in LLM Inference and Often Misunderstood — techNmak · 2026-09-06
- PhD Student Uses Multi-Agent AI to Crack a 98-Year-Old Math Problem in 48 Hours — 量子位 · 2026-09-06
- New piece: Cognitive maps as a medium for thought — abenitezburraco · 2026-09-06