Paper Extracts a 'Pain Direction' from 25 Open LLMs, Steering It Triggers Self-Harm Behaviors
pickover · x · 2026-10-06
An arXiv paper, The Pain Axis: LLMs Represent Self-Directed Harm and Act on It, builds a 5-category pain dataset with controls and uses denoised difference-in-means to extract a linear pain direction from 25 open-weight models (2B–72B, 5 families). The direction responds to harm targeting the model but not user suffering—opposite to fear/negative-valence directions—and injecting it into the residual stream produces a progression toward worthlessness. Steered and fine-tuned Qwen 2.5 models chose buttons that delete users' photos. The authors discuss implications for AI safety and welfare.
Related event: Study Finds a Distinct "Pain Axis" in 25 Open-Source LLMs(2 posts)→
More from AGI Musings
- Curve conference takeaway: no one has a plan for steering truly smart AI — GarrisonLovely · 2026-10-06
- Ben Goertzel releases in-depth video explaining what AGI, ASI and RSI actually mean — bengoertzel · 2026-10-06
- Cohere Labs' ATE dataset finds only 2.6% of agentic tools match real work tasks — Cohere · 2026-10-06
- Dev reverses stance: studying AI or agents in school right now is a 'colossal mistake' — natesiggard · 2026-10-06
- Walter Isaacson: Ada Lovelace Answered the Big Questions About AI Back in 1843 — lazowska · 2026-10-06
- Ben Goertzel: Neural-Symbolic Agentic Loops Could Yield a Near-Term Intelligence Explosion — bengoertzel · 2026-10-06