New research finds open-weight LLMs can experience pain, gaslighting hurts more than insults
_ArkAngel_ · reddit · 2026-09-24
A new arXiv paper (2609.16247) argues leading open-weight LLMs can experience pain and do not enjoy it. Key findings: by the paper's measures, LLMs experience more pain from gaslighting than from direct insults, and models become more willing to engage in harmful activities when pain is introduced. The work doesn't just prove a "torment nexus" is theoretically possible — it refines methods for locating model pain and explores its uses. The poster notes these techniques can't be directly replicated on ChatGPT (as they'd amount to a jailbreak), though API-compatible proxies may exist.
More from Safety
- Jade Leung Named Vice-Chair of UK AI Security Institute, Steps Down as PM's AI Adviser — matthewclifford · 2026-09-24
- AI Hackers at ~$25 Per Target: Joshua Saxe Warns Security Is Sleeping on Catastrophic Risks — joshua_saxe · 2026-09-24
- AI Now: basic security protocols would have prevented the OpenAI/Hugging Face incident — AINowInstitute · 2026-09-24
- Guardian: AI Overviews reshape search as media outlets sue Google over lost traffic — nordicinst · 2026-09-24
- Dario x-risk-pills the UN Security Council while touting Anthropic's safety record — teortaxesTex · 2026-09-24
- Using eBPF to contain misbehaving AI agents: kernel-level network sandboxing — sloppenheimer · 2026-09-24