Do LLMs "Feel"? Emotion Circuits Discovery and Control
Chenxi Wang, Yixuan Zhang, Ruiji Yu, Yufei Zheng, Lang Gao, Zirui Song, Zixiang Xu, Gus Xia, Huishuai Zhang, Dongyan Zhao, Xiuying Chen
cs.CL, cs.AI
2025-10-13
Maps emotion circuits in Llama-3.2-3B; modulating a few neurons and attention heads drives emotion expression to 99.65% accuracy, beating prompting and steering.
People already lean on chatbots like GPT-4o for emotional support, yet nobody can explain where an LLM's emotional output comes from, let alone control it at the mechanism level. Prior findings stop at signals: activations do encode emotion, and Tigges et al. (2024) showed emotions behave as roughly linear directions in activation space. Control methods such as steering vectors and prompt engineering work in practice but explain nothing. Deeper attempts stalled too: Tak et al. (2025) found layer-wise repairs unstable across layers, and the "emotion neurons" of Lee et al. (2025) barely changed output when masked.
The paper splits the question into three: do LLMs hold context-agnostic emotion mechanisms, what form do they take, and can they support universal emotion control?
Main experiments run on Llama-3.2-3B-Instruct, with the full pipeline reproduced on Qwen2.5-7B-Instruct.
Causal validation is strict. Ablation shows necessity: scores crash at k = 2 and 4, then barely move as the neuron count grows 30x. Enhancement shows sufficiency: injecting emotion difference vectors into top components evokes the emotion with no emotion instruction in the prompt. Random components given the same interventions do almost nothing, averaged over 10 seeds.
| Method | Emotion-expression accuracy (held-out test set) | Surprise |
| Prompting | 98.96% | not broken out in the main text |
| Steering vectors (layers 11-20) | 91.22% | 67.71% |
| Circuit modulation | 99.65% | 100% |
Circuit modulation injects the emotion difference vectors (λ = 1.0) into the components the circuit selected, with no emotion in the prompt and greedy decoding. At the representation level, hidden states split into emotion clusters from layer 9 onward; by layer 12 anger sits next to disgust and sadness next to fear, matching human affective intuition. Circuits for different emotions share almost no neurons (overlap 0.056 ± 0.033) but about half their attention heads (0.454 ± 0.047), a dual architecture of emotion-specific MLP subcircuits plus shared attention pathways.
For interpretability work, this is a full workflow that moves from linear directions to component-level circuits, with reusable protocols at every step and the code released. For generation control, it is one level finer than steering: instead of adding a single global vector to the residual stream, it modulates specific neurons and heads, which lifts surprise from 67.71% under steering to 100%. The output reads naturally too; exclamations like "Whoa?!" surface without any instruction. Companion products, emotional-support dialogue, and damping a model's emotional output for safety are the direct use cases. The boundary is equally clear: validation covers 3B and 7B open models, a solid step in mechanistic interpretability that does not license extrapolation to frontier scale.
The authors list three: English-only inputs, only the six Ekman basic emotions, and no test of circuit stability under fine-tuning or transfer. A close read adds caveats. Success is labeled by GPT-4o-mini (with manual consistency checks by the authors), an LLM judging a subjective property, so 99.65% may run optimistic. The "independent" test set matches SEV's distribution and size, which stretches the out-of-domain claim. And the "Feel" in the title is a hook: every result concerns the machinery of emotional expression, not subjective experience, so the paper cannot ground claims about whether AI can feel anything.