36 of Gemma's 42 layers match human valence, and the V is nearly linear

Latent Structure of Affective Representations in Large Language Models

Benjamin J. Choi, Melanie Weber

cs.LG, cs.AI

2026-04-08

Gemma-2-9B, Mistral-7B, and LLaMA-3-70B hit 13.7–21.3% zero-shot on GoEmotions, yet mid-layer activations form a V aligned with human valence and nearly linear.

What problem this solves

Activation-geometry papers rarely have a map worth checking against. A blob in a projection can be a real concept layout, or a shape the embedding method invented. Emotion has a map from psychology: valence, pleasant to unpleasant, and arousal, calm to excited. Older charts draw a ring around neutral. Later work draws a V, because more extreme valence tends to come with higher arousal. Discrete labels and continuous axes occupy the same space, so a claimed alignment is easier to audit than a generic topological statistic.

The linear representation hypothesis treats a high-level concept as a direction, so a linear probe can read it and a nudge along that direction can change the output. The manifold hypothesis says the structure is curved. Earlier emotion studies mostly score questionnaire answers or output probabilities. This paper looks at hidden states after a zero-shot read of emotional text: do the coordinates sit on the human valence map, how bent are they, and do linear tools still work.

Method

The corpus is GoEmotions, about 58,000 English Reddit comments with 27 emotions plus Neutral, kept only when a single label has rater agreement (about 83%). Gemma-2-9B (42 layers), Mistral-7B (32 layers), and LLaMA-3-70B-Instruct (80 layers) classify zero-shot, with no fine-tuning on this set. Qwen models and some other LLaMA checkpoints miss the recognition bar and stay out of the main analysis.

Each example becomes one mean-pooled vector per layer. Geometry is fit only on correct classifications. Correct counts run from about 100 to 2,200 on Gemma and about 100 to 1,300 on Mistral, so a joint 28-way classifier would be owned by the head classes. Comparisons are pairwise: downsample the majority emotion, fit L2 logistic regression, and use 80:20 test accuracy as a dissimilarity. An emotion needs at least 100 correct calls to enter, which leaves 20 classes for Gemma and 16 for Mistral. Cosine distance leaves the main picture in place.

Classical MDS turns the dissimilarities into Euclidean coordinates. Isomap swaps straight-line distance for shortest paths on a k-nearest-neighbor graph, then runs MDS, which unrolls curvature. The human reference is ANEW, normative valence and arousal ratings; 17 of the 28 classes have scores there. After centering, a scaled orthogonal Procrustes fit (one rotation, one uniform scale) is tested against 2,000 label shuffles.

Misclassified activations never enter probe training. They do not sit on the predicted class. They fall between the gold label and the label the model emitted, closer to the separating hyperplane than correct points are. Early layers lean toward the emotion in the prompt; later layers lean toward the wrong label. Distances to those hyperplanes, across layers, train a second logistic regression, calibrated on a validation split into a probability that the call was correct. A pair needs at least 25 correct and 25 incorrect examples, split 60:20:20.

Results

Zero-shot accuracy is 19.4% for Gemma-2-9B, 13.7% for Mistral-7B, and 21.3% for LLaMA-3-70B-Instruct. Chance on 28 classes is about 3.6%. The best model is right about one time in five.

On items the model gets right, pairwise separability is much higher. Gemma stays above 0.9 mean test accuracy in every layer. Mistral moves from about 0.55 to about 0.89, highest in middle layers and a bit lower at the end. In the 2D MDS plots (Gemma layer 20, Mistral layer 27), positive emotions form one arm, negative emotions the other, and neutral sits near the vertex.

ModelZero-shotSignificant ANEW layersUncertainty accuracyMajority baselineECE
Gemma-2-9B19.4%36/4277.6% (AUC 0.813)69.4%0.011
Mistral-7B13.7%17/3285.7% (AUC 0.871)82.2%0.007
LLaMA-3-70B-Instruct21.3%34/8080.1% (AUC 0.822)75.6%0.011

Mistral's 17 significant layers include its last 14. The 34 significant layers of the 70B model also sit late. Early Mistral layers barely separate emotions and show no V.

The spectrum is spread out. Participation ratios, a count of how many directions carry the variance, are about 17 for Gemma and about 9 to 14 for Mistral. That is not a low-dimensional manifold. Gemma has a significant first eigengap in 16 of 42 layers, a recurring valence-like axis, not a fixed intrinsic dimension. Isomap beats classical MDS on neighborhood trustworthiness almost only at rank 1, with median gains of 0.161 and 0.127. At higher rank the mean gap sits near zero. In one dimension Isomap unrolls the V into a line; in two dimensions the arms are already visible, so unrolling adds nothing. Geodesic-to-Euclidean distance ratios, 10th to 90th percentile, run from 1.00 to 1.80 on Gemma and from 0.80 to 1.42 on Mistral. The bend is mild. Linear coordinates still work.

The uncertainty model covers 180 pairs on Gemma (3,858 correct vs 8,748 wrong) and reaches 77.6% against a 69.4% majority baseline. Mistral, on 122 pairs (5,254 vs 15,216), reaches 85.7% against 82.2%. LLaMA-3-70B-Instruct reaches 80.1% against 75.6%. AUCs are 0.813, 0.871, and 0.822. Expected calibration errors are 0.011, 0.007, and 0.011.

Why it matters

Linear probes assume a concept is a direction. The geometry measured here is a parabola globally, and close to Euclidean once two dimensions are allowed. Isomap pays off when the V is forced onto a line, and barely otherwise. Linear read and write of affect is a local approximation with measurements behind it. An appendix reports that probe directions can shift the valence of generated text. The main text gives no effect size.

Distance to the hyperplane estimates whether an emotion call was correct. AUCs of 0.813 to 0.871 and calibration errors near 0.01 mean the ranking and the probabilities are usable. Accuracy rises only 3.5 to 8.2 points over the majority class, and the task is that one classification, not open-ended hallucination.

Lining up with the human valence chart means the coordinates match. It does not mean the model feels anything. A similar parabola in human EEG is mentioned from an appendix, with no numbers in the main text. What you can take away is a check for likely misclassification, and an affect direction you can intervene on.

Limitations

The paper lists three gaps. Other model families, and what emotion-specific fine-tuning does to the geometry, are untested. Pairwise regression plus MDS and Isomap can miss finer structure. GoEmotions is English Reddit, so culture and language skew the picture, and attention was not split from the MLP. The dissimilarity is also not a true distance: a layer sometimes shows one or two negative eigenvalues, under 1% of the variance.

The V is drawn on correct items only. Zero-shot accuracy is 13.7% to 21.3%, so most calls are wrong, and those activations lie between the two labels. Treating the correct subset as proof that the model learned human emotion geometry overreaches. Qwen and some LLaMA models were dropped for missing the recognition bar, and that selection is not measured on its own. The alignment test uses the 17 classes that overlap ANEW. Mistral's 85.7% uncertainty accuracy stands next to an 82.2% majority baseline.

Terms

Source

What people are saying

Related papers

All paper explainers