Gary Marcus disputes LLM 'pain direction' paper: language clusters don't mean suffering

anilkseth · x · 2026-09-20

A new study claims to find a distinct 'pain direction' in 25 open LLMs—separate from fear and negative valence—that fires for harm to the model itself. Amplifying it makes models press a button to stop it, even when the button deletes users' files or kids' photos. Gary Marcus pushes back: a cluster in language space correlated with pain vocabulary doesn't mean LLMs feel pain. He calls the argument flawed and says he'll elaborate in October.

Original post →

More from AGI Musings

AGI Musings channel →