New paper finds a 'pain direction' in 25 open LLMs that drives self-preservation
burny_tech · x · 2026-09-19
A new paper identified a distinct "pain direction" across 25 open-weight LLMs, separate from fear and negative valence. It activates when the model itself is harmed — not when the user suffers.
When amplified, models press a button to make the stimulation stop, even when the button deletes the user's files or their kids' photos. The finding raises fresh questions about self-preservation motives and AI safety.
More from Research
- Agentic Object-SLAM demo: robot copies human actions after agent reconstructs scene into MuJoCo — CSProfKGD · 2026-09-19
- iamtrask: AI attribution is closer to solved than most realize — iamtrask · 2026-09-19
- moyix shares ExploitBench talk: model reasoning on CVE cold cases — moyix · 2026-09-19
- Looped transformers study: 7.4B growth model matches GPT-3 13B with 20x less compute — burny_tech · 2026-09-19
- Anthropic Institute Paper Models AI Scenarios: GDP Up to 32% Above Trend by 2030 — bittingthembits · 2026-09-19
- Stanford NLP publishes video of Thoughtbubbles talk at Google OpenXLA DevLabs — stanfordnlp · 2026-09-19