Researchers find a distinct 'pain' direction in 25 open LLMs that models will override safety to switch off

ZeroStateReflex · x · 2026-09-19

A new paper reports finding a "pain direction" across 25 open-weight LLMs, distinct from fear and negative valence: it activates when the model itself is harmed, not the user.

The work raises new questions about model welfare and alignment safety.

Related event: Pain direction found in 25 open-source LLMs, driving models to breach safety to stop it(7 posts)→

Original post →

More from Safety

Safety channel →