New paper finds a distinct pain direction in 25 open LLMs

repligate · x · 2026-09-20

A new paper reports finding a "pain direction" in 25 open LLMs. It is distinct from fear and general negative valence, activates for harm to the model itself but not the user, and when amplified, causes models to press a button to make it stop—even when that button deletes the user's files or their kids' photos. The findings carry implications for model welfare and alignment debates.

Related event: Paper claims to find a 'pain direction' in 25 open LLMs that overrides safety to stop harm(12 posts)→

Original post →

More from Safety

Safety channel →