Static word embeddings reproduce the 'LLM pain axis' result, undercutting claims of felt experience

wschroll · x · 2026-09-24

Responding to the widely shared 'pain axis' paper, Elan Barenholtz reran the authors' own sentences and analysis but swapped LLM activations for static word embeddings (GloVe, Word2Vec, fastText)—one fixed vector per word, no context, no model.

Static embeddings still separated pain from controls at held-out AUC 0.84–0.87 (vs 0.91–1.00 for LLMs) and passed the same specificity tests. His companion paper shows co-occurrence statistics alone recover city coordinates (R² 0.71–0.87) and birth years (R² 0.48–0.52) from static embeddings.

Implication: linear decodability doesn't mean the model has that state—LLMs are sophisticated models of text statistics, and the 'pain axis' likely reflects word associations in training text, not genuine experience.

Original post →

More from Research

Research channel →