Circuit interpretability hits a wall as ablations miss across tasks

A COLM 2026 workshop paper finds extracted circuits lack specificity, with ablations harming other tasks; researchers explain why component-level specificity is low while neuron-level is higher, citing related SAE work on concept manifolds.

2026-09-05 ~ 2026-09-05 · 4 related posts