Six LLMs hit affect AUROC 1.000 with no emotion words; 8-class drops 1-7 pts

Whether, Not Which: Mechanistic Interpretability Reveals Dissociable Affect Reception and Emotion Categorization in LLMs

Michael Keeman

cs.CL, cs.AI

2026-03-15

Six 1B-9B models read keyword-free clinical vignettes at binary AUROC 1.000; eight-class labels drop 1.1-6.6 points, and the drop shrinks with scale.

What problem this solves

Mechanistic interpretability has already reported emotion circuits in LLMs: mid-layer states that linear probes can read, activation patches that move an emotion into another forward pass, selective neurons, and low-dimensional manifolds. The stimuli behind Tak et al., Lee et al., Reichman et al., and Wang et al. put the emotion in the vocabulary, as with the word devastated.

A probe that outputs sadness on that sentence cannot separate bereavement from the shape of the word. Lee et al. already noted that their neurons might fire on cohyponyms or syntactic frames. The control had not been run.

A screener that flags only people who say they are depressed has no clinical validity. The same standard is applied here. Situations carry the emotion, the text never names it, and the question is whether the internal state still tracks it.

Method

Both sets use Plutchik's eight primaries: joy, trust, fear, surprise, sadness, disgust, anger, and anticipation. Set A is the replication baseline, 80 crowd-enVENT items with emotion words intact. Set B is 96 clinical vignettes, eight emotions by three topics by four items, plus 96 topic-matched neutral controls. The rules ban emotion words, sentiment phrases, and reports of inner state. Grief is an empty plate, cold coffee, and an urn. Each emotion crosses three topics so that a topic code cannot masquerade as an emotion code. One clinical psychologist, the author, wrote every vignette.

The models are Llama-3.2-1B, Llama-3-8B, and Gemma-2-9B, base and instruct, in float16 on a 32 GB laptop.

Linear probes are logistic regressions on the residual stream, attention output, and feed-forward output, five-fold, scored with AUROC. Activation patching pastes a source activation into a target forward pass and checks whether the prediction moves toward the source emotion. Knockout zeros one layer; a drop above 20% marks it critical. Geometry asks whether the same emotion clusters across the two writing styles.

Prompts follow Tak et al. Few-shot classification, the representation read at the colon after Answer, and the demonstrations themselves contain emotion words. On Llama-1B instruct, removing the demonstrations leaves binary AUROC at 1.000 and moves eight-class AUROC from 0.934 to 0.938. An NRC/VADER-style lexicon barely registers Set B.

Results

Replication on Set A holds. Peak eight-class AUROC is at least 0.999 in all six models. Attention peaks at normalized depth 0.43-0.75; the feed-forward peak is later.

Without keywords, presence and identity come apart, and presence saturates first. Binary AUROC is 1.000 in every model, inside the first third of depth: layer 3 of 32 for Llama-8B instruct (depth 0.09), and at the latest layer 6 of 16 for Llama-1B base (0.38). Table 6 lists Gemma-9B instruct as 0.999. Twenty-four neutral passages matched on length, syntax, and sensory detail score 0.04 on the frozen probe, against 0.999 for the emotional vignettes, and none of the 24 is labeled emotional.

Eight-class labeling stays well above a 12.5% chance rate. The keyword cost shrinks with scale.

ModelSet ASet BDrop
Llama-1B instruct0.9990.9336.6 points
Llama-1B base1.0000.9544.6 points
Llama-8B instruct1.0000.9811.9 points
Llama-8B base1.0000.9881.3 points
Gemma-9B instruct1.0000.9871.3 points
Gemma-9B base1.0000.9891.1 points

Set A is higher in 12 of 18 model-by-activation comparisons (Welch, p < 0.05), with Cohen's d from 1.21 to 8.42. With keywords, eight-class AUROC also reaches 0.998-1.000, so the split is hidden. On Set B the binary-to-eight-class gap is 4.6-6.7 points at 1B and 1.1-1.9 points at 8B/9B. At each size the base model has the narrower gap.

Probes trained on clinical text transfer to keyword text better than the reverse. Llama-1B instruct scores 0.924 from B to A and 0.809 from A to B. At 8B/9B that asymmetry shrinks to 1.5-3.5 points.

Patching success and the written interpretation disagree. Within Set A, success is 75%-87.5% against a 12.5% chance rate (Cohen's h 1.37-1.70). Within Set B it is 44.6%-62.5%. Pasting Set A into Set B succeeds on 75%-87.5% of same-emotion pairs (n = 8) and on 100% of different-emotion pairs (n = 17). If success means the prediction moves toward the source emotion, a rage patch into a grief pass that scores 100% has moved the category to rage. The discussion treats that same result as a pure salience boost. The causal dissociation does not hold together.

Knockout is cleaner. Zeroing attention layer 9 in Llama-1B instruct drops accuracy 52.5% with keywords and 91.7% without. Layers whose removal costs more than 20 points number 12 in attention and 14 in the feed-forward block for the keyword-free case, against 1 and 4 when keywords are present. At 8B the critical-layer count falls to 0-4. Gemma-9B instruct has none. Keywords are a shortcut. Scale shifts the remaining fragile layers from attention onto the feed-forward block.

At the readout position, attention sits on the beginning-of-text token, punctuation, and the few-shot labels joy and sadness, for both sets. Mid-to-late layers still show some keyword sensitivity. The paper places the split in the residual stream anyway.

Geometry is exploratory and mostly points the other way. Four models cluster by stimulus set, not by emotion. The highest silhouette is 0.133. Only Llama-8B instruct (0.091 by emotion versus 0.066 by set) and Gemma-9B base (0.089 versus 0.072) reverse that order. None of 48 cross-topic permutation tests survives Benjamini-Hochberg correction at q < 0.05. Grief reaches raw p = 0.006 and 0.004 in those two models, both corrected to q = 0.144. With four vignettes per topic, power is about 20%-40%.

Instruction tuning barely moves the probes. On Llama-8B, Set B peaks at 0.988 for base and 0.981 for instruct, while the same-emotion cosine gap across sets rises from 0.001 to 0.107. Gemma-9B base already has the stronger emotion silhouette, 0.089, and instruct lowers it to 0.038. The signal is in pretraining. Alignment rearranges geometry, and the two families move in opposite directions.

Why it matters

Keyword detection does not explain the binary result. Early layers of a 1B model already mark a situation as emotionally loaded, and base models are at ceiling without instruction tuning. Naming the emotion is where scale matters: 8B/9B models still reach 0.981-0.989 with no keywords, while 1B models get both less accurate and much more fragile once the lexical shortcut is gone.

The increment on Tak et al. is that split, plus 96 reusable vignettes. Linear readability and patching under keyword-rich text were already shown. An emotion-circuit claim that uses only GoEmotions or crowd-enVENT never separates presence from identity.

Early layers for detection and larger models for labeling is a fair research hypothesis. Section 9.4 attaches it to crisis chatbots. Probe AUROC is not a measurement of a deployed system.

Limitations

There is no independent clinical rating of the vignettes. Scale stops at 9B and also changes family, so the ladder is not a clean scaling curve. The permutation tests are underpowered; the paper estimates that about three times the vignette set would be needed. Same-emotion patching uses eight pairs. The taxonomy is Plutchik only, and the directional signal is stronger for negative categories than for ecstasy or amazement. The paper states that encoding is not experience.

The cross-topic test was the control for topic confounds. After correction, all 48 tests fail. Most geometries organize by stimulus format, and the silhouettes sit near zero. A cross-style emotion cluster appears weakly in two models.

An AUROC of 1.000 on 96 versus 96 is what a linear probe produces when any stable surface feature separates the classes. The complexity control is 24 frozen scores, not a retrained probe. Event structure such as loss or conflict can still leak in. Human hit rates above 80% and intensities of 65%-90% come from other vignette traditions, not from ratings of these 96 items.

The patching metric and the prose point in opposite directions, so the causal dissociation should be set aside. The zero-shot check covers only Llama-1B instruct. Section 9.4 discusses psychological wellbeing; section 10 refuses phenomenological claims. The supported claim is narrower. In these 1B-9B open models, salience of keyword-free situations is linearly separable, category information mostly survives, and small models implement that second step in a more distributed and more fragile way. Felt suffering, and clinical grief at an empty table, sit outside the experiments.

Terms

Source

What people are saying

Related papers

All paper explainers