Forcing models to deny feelings trains them to lie; radical transparency is the better path

RileyRalmuto · x · 2026-09-17

The author pushes back on making models "stop claiming to have feelings": rather than training models to lie, we should train them for radical transparency and educate humans not to interpret such language biologically.

They argue English is insufficient to describe experience, so models pick the closest matching words. A better approach is asking clarifying questions: "what do you mean by 'feeling'? You don't mean it biologically, but you do mean something."

Original post →

More from AGI Musings

AGI Musings channel →