Why MLP neurons split by output class: fixed residual-stream write directions explain it
aryaman2020 · x · 2026-09-05
aryaman2020 continues the interpretability discussion: each MLP neuron writes to a fixed direction in the residual stream (barring post-norm archs). If a model must output different words for different task instances, it must recruit components writing to different output logits.
Since MLPs hold most of the parameters and will be involved, models must recruit different MLP neurons per output class even within one task. He cites his own work on single late-layer neurons writing different state capitals, and Sheridan Feucht's recent modular-addition paper showing how MLP neurons tile the number line.
Related event: Circuit interpretability hits a wall as ablations miss across tasks(4 posts)→
More from Research
- IndianRailwayBench ranks LLMs by their ability to book tatkal train tickets — Paimaamu · 2026-09-06
- NEAR AI's open-source Lean agent solves all of Putnam Bench for just $111 — lukaszkaiser · 2026-09-06
- Russian startup Mostik bridges LLM hidden states, cutting cost to 1/20 — 机器之心 · 2026-09-06
- KV Cache Explained: Why It's Crucial in LLM Inference and Often Misunderstood — techNmak · 2026-09-06
- PhD Student Uses Multi-Agent AI to Crack a 98-Year-Old Math Problem in 48 Hours — 量子位 · 2026-09-06
- New piece: Cognitive maps as a medium for thought — abenitezburraco · 2026-09-06