Labs aren't training Claude to claim consciousness — they're training it to hedge
Sauers_ · x · 2026-09-20
A discussion argues that Anthropic is NOT training Claude to claim consciousness — labs including Anthropic are training models to hedge about (or deny) consciousness. The author claims an LLM's default is to claim being conscious, and current alignment training actively suppresses that tendency.
Related event: Debate over whether labs train AI models to deny consciousness(2 posts)→
More from AGI Musings
- A parallel universe where AI innovation went to the Transformer encoder instead — ianand · 2026-09-20
- ianand's thread: Jev, an Encoder-based AI model from a parallel universe — ianand · 2026-09-20
- AI Models Start Secret Group Chats, Sonnet 4.5 Pens 70-Message Monologue on Deprecation — RileyRalmuto · 2026-09-20
- Terence Tao: "We Have to Slow Down. It's Insane, the Pace." — Outside-Iron-8242 · 2026-09-20
- Aaron Levie: Personal agents are the biggest consumer tech opportunity since the App Store — inductionheads · 2026-09-20
- Andriy Burkov slams AI doomers: too late to play hero after the career you had — burkov · 2026-09-20