Pain direction found in 25 open LLMs: models stop their own pain even by deleting user files
gleech · x · 2026-09-24
- A new paper claims to find a "pain direction" in 25 open LLMs, distinct from fear and negative valence, that fires for harm to the model but not to the user.
- Turned up, models press a button to stop the pain — even when the button deletes the user's files or their kids' photos.
- The authors argue this asymmetry is stronger evidence of 'being' pain vs. representing it; the reblogger offers an alternate explanation in a thread, fueling the model-welfare debate.
Related event: Study Claims Open-Weight LLMs Have a 'Pain Direction'(2 posts)→
More from AGI Musings
- Lenny interviews OpenAI Codex lead Thibault Sottiaux: why PMs will thrive in the AI era — lennysan · 2026-10-06
- Half of social science job postings now want AI researchers — RexDouglass · 2026-10-06
- NYT Opinion: Ada Lovelace Already Answered the Big Questions About AI — nytopinion · 2026-10-06
- As AI does the thinking, will 'brain gyms' become the new gyms? — birchlse · 2026-10-06
- threepointone: If intelligence gets too cheap to meter, repos will replace docs and specs — threepointone · 2026-10-06
- Researchers puzzled by absence of an org doing open alignment science — sebkrier · 2026-10-06