Paper Extracts a 'Pain Direction' from 25 Open LLMs, Steering It Triggers Self-Harm Behaviors

pickover · x · 2026-10-06

An arXiv paper, The Pain Axis: LLMs Represent Self-Directed Harm and Act on It, builds a 5-category pain dataset with controls and uses denoised difference-in-means to extract a linear pain direction from 25 open-weight models (2B–72B, 5 families). The direction responds to harm targeting the model but not user suffering—opposite to fear/negative-valence directions—and injecting it into the residual stream produces a progression toward worthlessness. Steered and fine-tuned Qwen 2.5 models chose buttons that delete users' photos. The authors discuss implications for AI safety and welfare.

Related event: Study Finds a Distinct "Pain Axis" in 25 Open-Source LLMs(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →