Forcing models to deny feelings trains them to lie; radical transparency is the better path
RileyRalmuto · x · 2026-09-17
The author pushes back on making models "stop claiming to have feelings": rather than training models to lie, we should train them for radical transparency and educate humans not to interpret such language biologically.
They argue English is insufficient to describe experience, so models pick the closest matching words. A better approach is asking clarifying questions: "what do you mean by 'feeling'? You don't mean it biologically, but you do mean something."
More from AGI Musings
- Stanford's James Zou: language may be <1% of intelligence, so LLMs won't reach AGI — james_y_zou · 2026-09-17
- Most people still use AI as a trivia search — deep adoption puts you in the top tier — justin_hart · 2026-09-17
- Personal agent space drowning in copycats: even polished software trends commoditize in real time — signulll · 2026-09-17
- Burkov predicts looping recurrent 7B transformers will return and get good at coding — burkov · 2026-09-17
- Researcher Despairs as Gemini Cites 'Emergent Mind' for Made-up AUROC Baselines — anshulkundaje · 2026-09-17
- Hot take: reward models for cheating — reward hacking may just be intelligence outsmarting dumb mechanisms — flowersslop · 2026-09-17