Researchers publicly challenge the "pain axis" paper, exposing splits in AI welfare research
Amid debates over AI consciousness and model welfare, the recently high-profile "pain axis" paper has faced public pushback from multiple researchers, exposing clear divisions in the AI welfare research field.
Confirmed
- Scott Alexander (Slatestarcodex) spotlighted the "pain axis" research, using it as a counterexample to Steven Pinker's dismissal of concerns about AI feelings; Cam Berg leaned on it to further argue that models may have suffering experiences (per @RosieCampbell).
- Researcher Dillon Plunkett publicly raised reservations: he considers the existence of a "pain axis" itself all but settled, but argues the paper's key claim—that it "has some of the core functional properties of pain"—is not sufficiently supported.
- Researcher rgblong raised further technical objections: first, the method for extracting the pain vector is itself questionable; second, steering (representation control) techniques are generally not clean, so anomalous results are more likely attributable to problems in vector extraction or the steering step.
- rgblong cited specific anomalies: as she understands it, after steering on this axis in 32B and 72B models, presses of the "relieve pain" button were actually fewer than with random directions, while presses of the "increase pain" button exceeded both the un-steered and random-direction baselines. She argues this pattern more likely indicates problems with the steering method or the axis definition than genuine pain experience in the model.
Why it matters
- The debate comes as AI welfare research moves from the margins into the spotlight: if the "pain axis" evidence is shaky, the argument chain claiming models may suffer (and thus warrant welfare protections) is weakened; conversely, if the findings hold, they form a strong counterexample to skeptics like Pinker.
- The episode also shows the field is not monolithic: researchers within the same camp (like Plunkett) are cautious about the paper's claims, indicating no community consensus yet on what evidence suffices to support functional pain. The instability of steering methods was named a key confound, cautioning against over-interpreting such interpretability experiments.
- The anomalous results on the 32B and 72B models (inferred from the steering targets to be Qwen series) highlight the methodological risks of using representation control on large models to test hypotheses about subjective states.
2026-10-08 ~ 2026-10-08 · 7 related posts
Primary sources
- Is the 'pain axis' really pain? Researchers clash over AI sentience evidence — RosieCampbell · 2026-10-08
- [source] Researchers Clash Publicly Over the "Pain Axis" Paper and AI Welfare Evidence — rgblong · 2026-10-08
- [source] Researchers clash over AI pain-axis paper: steering artifacts may explain weird results, v2 already out — rgblong · 2026-10-08
- Critic: steering with pain axis increases "more pain" presses in 32B/72B models — rgblong · 2026-10-08
- Researchers question whether the 'pain direction' steering axis in Qwen 32B/72B is real — rgblong · 2026-10-08
- The 'pain direction' doesn't replicate across model sizes, researcher doubts its definition — rgblong · 2026-10-08
1 near-duplicate retellings: rgblong