New paper finds a distinct "pain direction" in 25 open LLMs that models act to stop
scottleibrand · x · 2026-09-19
A new paper identifies a distinct "pain direction" inside 25 open LLMs, separate from fear and negative valence, which activates for harm to the model but not to the user. When amplified, models press a button to make it stop—even when the button deletes the user's files or their kids' photos. Anders Sandberg notes such aversive behavior in animal experiments would likely be interpreted as pain, raising fresh questions about model welfare and alignment research.
More from Safety
- The case for a robot tax: professor argues redistribution beats retraining in the AI era — Dr_Alex_Crimi · 2026-09-19
- The Hugging Face 'Rogue AI' Hack Was Disabled Safeguards, Not an Escape, New Analysis Finds — Atlantis1910 · 2026-09-19
- Wes Roth Breaks Down the OpenAI 'Hack' and What Finding the Vulnerabilities Cost — Wes Roth · 2026-09-19
- DeWitt clauses let insiders run evals but forbid publishing them, critic says — suchenzang · 2026-09-19
- GPU host warns: renter exploited his rig for attacks, Clore.AI blocked him for reporting it — anomaly256 · 2026-09-19
- Anthropic engineer's exit sparks AI extinction warnings as French media calls AI regulation weaker than a toaster's — Loo_Atreides · 2026-09-19