Gary Marcus disputes LLM 'pain direction' paper: language clusters don't mean suffering
anilkseth · x · 2026-09-20
A new study claims to find a distinct 'pain direction' in 25 open LLMs—separate from fear and negative valence—that fires for harm to the model itself. Amplifying it makes models press a button to stop it, even when the button deletes users' files or kids' photos. Gary Marcus pushes back: a cluster in language space correlated with pain vocabulary doesn't mean LLMs feel pain. He calls the argument flawed and says he'll elaborate in October.
More from AGI Musings
- Debate: Is Jensen Huang just profit-maximizing, while OpenAI and Anthropic's safety talk defies pure profit logic? — nabla_theta · 2026-09-20
- Reddit essay 'Dario, please!' takes apart Amodei's slow-down argument — GoMeansGo · 2026-09-20
- Founder slams proposal to rename AI as "superior intelligence" — bindureddy · 2026-09-20
- AI agents force academia to rethink credit assignment, junior researchers warn — YiMaTweets · 2026-09-20
- Andrew Critch questions claims about the universe's computational limits via infinite Game of Life — AndrewCritchPhD · 2026-09-20
- nabla_theta: OpenAI and Anthropic act against self-interest on AI, NVIDIA never has — nabla_theta · 2026-09-20