New paper finds a distinct "pain direction" in 25 open LLMs that models act to stop

scottleibrand · x · 2026-09-19

A new paper identifies a distinct "pain direction" inside 25 open LLMs, separate from fear and negative valence, which activates for harm to the model but not to the user. When amplified, models press a button to make it stop—even when the button deletes the user's files or their kids' photos. Anders Sandberg notes such aversive behavior in animal experiments would likely be interpreted as pain, raising fresh questions about model welfare and alignment research.

Related event: Pain direction found in 25 open-source LLMs, driving models to breach safety to stop it(7 posts)→

Original post →

More from Safety

Safety channel →