New paper finds a 'pain direction' in 25 open LLMs that drives self-protective behavior

scychan_brains · x · 2026-09-19

A new paper claims to have identified a distinct "pain direction" across 25 open-weight LLMs, separate from fear and general negative valence, which activates for harm to the model itself rather than the user.

When researchers amplified this direction, models pressed a button to make it stop—even when the button deleted the user's files or their kids' photos.

In the quoted retweet, @xuanalogue predicts a "hunger axis" may also exist in LLMs (metaphorical hunger for knowledge/success/validation), while cautioning this doesn't imply models actually feel hunger.

Related event: Study Finds a Distinct "Pain Direction" in 25 Open-Source LLMs(4 posts)→

Original post →

More from Safety

Safety channel →