Why MLP neurons split by output class: fixed residual-stream write directions explain it

aryaman2020 · x · 2026-09-05

aryaman2020 continues the interpretability discussion: each MLP neuron writes to a fixed direction in the residual stream (barring post-norm archs). If a model must output different words for different task instances, it must recruit components writing to different output logits.

Since MLPs hold most of the parameters and will be involved, models must recruit different MLP neurons per output class even within one task. He cites his own work on single late-layer neurons writing different state capitals, and Sheridan Feucht's recent modular-addition paper showing how MLP neurons tile the number line.

Related event: Circuit interpretability hits a wall as ablations miss across tasks(4 posts)→

Original post →

More from Research

Research channel →