Interp researcher: models' pain-axis 'lashing out' may be 20% moral injury, not pain

rgblong · x · 2026-10-08

An interpretability researcher questions the reading that a 'pain axis' causes models to lash out destructively. He suspects the effect may come from 20% of the axis being extracted from 'moral injury'—a concept that bakes in a heavy dose of 'I did a bad thing'—so the destructive behavior may stem from conflated concepts rather than pain itself.

Related event: Researchers Challenge the 'Pain Axis' Paper as AI Welfare Debate Deepens(19 posts)→

Original post →

More from Research

Research channel →