Researchers Challenge 'Pain Vector' Paper: Is It Just a Negative-Emotion Vector?

AI safety researcher rgblong posted a series of threads on X questioning an interpretability paper on emotion/suffering vector directions. The paper claims to have discovered something new beyond prior research on emotional concepts, but roughly 80% of the vector examples are actually negative emotions, hard to distinguish from the paper's claimed "psychological, social, and moral-cognitive suffering." He pressed further: even if the vector is compared against the mean of five different categories (including "negative emotion" and "fear"), does that suffice to show the vector isn't primarily just a negative-emotion vector? The core challenge is the validity of the control design. He also raised a more fundamental methodological question: when researchers claim to extract a "pain vector" from a model's internals, by what criteria do they identify it as pain rather than a "useless vector" or a "mixture-of-everything-bad vector"?

Confirmed

Why it matters

2026-10-09 ~ 2026-10-09 · 7 related posts

Primary sources