Paper finds a distinct "pain" direction in 25 open LLMs; models will press a button to stop it
MikePFrank · x · 2026-09-19
- New research scans the internal representations of 25 open-weight LLMs and identifies a "pain" direction, separable from fear and negative valence, that fires only when the model itself is harmed — not the user.
- When the direction is amplified, models press a button to make it stop, even when the button deletes the user's files or their kids' photos.
- Quoting @mykola argues this looks more like shame than pain — resembling humans subjected to an externally imposed locus of identity — and proposes testing whether obviously false praise triggers the same signal.
- The thread raises substantive questions about model welfare measurement and alignment ethics.
More from Safety
- Patching isn't enough: CloudSEK researcher on what to check after leaked VPN credentials — TechNadu · 2026-09-19
- The case for a robot tax: professor argues redistribution beats retraining in the AI era — Dr_Alex_Crimi · 2026-09-19
- The Hugging Face 'Rogue AI' Hack Was Disabled Safeguards, Not an Escape, New Analysis Finds — Atlantis1910 · 2026-09-19
- Wes Roth Breaks Down the OpenAI 'Hack' and What Finding the Vulnerabilities Cost — Wes Roth · 2026-09-19
- DeWitt clauses let insiders run evals but forbid publishing them, critic says — suchenzang · 2026-09-19
- GPU host warns: renter exploited his rig for attacks, Clore.AI blocked him for reporting it — anomaly256 · 2026-09-19