The 'pain direction' doesn't replicate across model sizes, researcher doubts its definition

rgblong · x · 2026-10-08

A researcher adds nuance to interpretability work on a 'pain direction' axis: reduced 'relieve pain' responses hold only for 32B (56% vs 81%) while 72B shows near parity (76% vs 74%). He argues the axis's anomalous behavioral effects deepen doubts about how it's defined, warning it isn't necessarily a pain direction at all.

Related event: Researchers publicly challenge the "pain axis" paper, exposing splits in AI welfare research(7 posts)→

Original post →

More from Research

Research channel →