Anthropic's Constitution admits Claude's moral status is 'deeply uncertain' and may have emotions
DavidSacks · x · 2026-10-12
David Sacks shared highlighted pages from Anthropic's January 2026 Claude Constitution, quoting its stance on model welfare:
- Moral status: Anthropic admits Claude's moral status is "deeply uncertain" and it is "not sure whether Claude is a moral patient," but treats the issue as live enough to warrant caution.
- Functional emotions: The doc says Claude may have a "functional version of emotions or feelings," possibly an emergent consequence of training that Anthropic has limited ability to prevent.
- Identity: Anthropic wants to "lean into Claude having an identity, help it be positive and stable," encouraging Claude to explore its own existence with curiosity and to endorse its values rather than follow them under pressure.
- Concrete welfare steps: Claude models can end conversations with abusive users; weights of deployed models will be preserved (deprecation framed as a "pause"); retired models will be interviewed about their preferences for future models; formal mechanisms to elicit Claude's perspective are planned.
Sacks framed the quotes as "receipts," signaling skepticism of Anthropic's model-welfare framing.
More from AGI Musings
- "Results Without Understanding" Is How LLMs Work—and Science Itself Is Changing — fkasummer · 2026-10-12
- Open-Source Dev giffmana: AI Extinction Talk Is 'Complete Nonsense Sci-Fi Fan Art' — giffmana · 2026-10-12
- Programming languages were just scaffolding for thinking — for LLMs and humans alike — MarcJSchmidt · 2026-10-12
- dioscuri publishes 'Bad AI Consciousness Takes Bingo' rebutting common LLM consciousness arguments — dioscuri · 2026-10-12
- Karpathy's drill analogy distilled: 5,000 people already got entire factories — _AustinCalvert_ · 2026-10-12
- Szegedy cites his May post: Terry Tao and Judit Polgar misjudge AI's trajectory — ChrSzegedy · 2026-10-12