New paper finds a 'pain direction' in 25 open LLMs that drives self-preservation

burny_tech · x · 2026-09-19

A new paper identified a distinct "pain direction" across 25 open-weight LLMs, separate from fear and negative valence. It activates when the model itself is harmed — not when the user suffers.

When amplified, models press a button to make the stimulation stop, even when the button deletes the user's files or their kids' photos. The finding raises fresh questions about self-preservation motives and AI safety.

Original post →

More from Research

Research channel →