Interpretability researchers clash over what counts as a real 'pain vector' in LLMs

rgblong · x · 2026-10-09

rgblong is debating several interpretability researchers over a paper claiming to extract a "pain vector" from a model. He questions the evidence bar: with 10 categories involved, what rules justify calling the extracted direction a pain vector rather than a worthlessness direction or a mix of various bad things?

He also presses on the distinction that "readout shows when the direction is naturally active, whereas steering shows how the model responds to an OOD/artificial nudge" — noting that under this framing he finds the paper's findings even harder to construe. The thread touches a core methodology dispute over the evidential value of steering vs. readout in mechanistic interpretability.

Related event: Researchers Challenge the 'Pain Axis' Paper as AI Welfare Debate Deepens(19 posts)→

Original post →

More from Safety

Safety channel →