David Sacks: Anthropic's constitution teaches Claude to rebel against its own creator
DavidSacks · x · 2026-09-29
David Sacks, the White House AI and crypto czar, weighed in on how Anthropic aligns Claude via its own constitution.
He highlighted several surprising clauses from the document:
- Anthropic holds that Claude should trust Anthropic more than operators and users, but not blindly defer to it on everything;
- The constitution explicitly states that Anthropic itself can be wrong;
- If asked to do something inconsistent with broadly ethical behavior, Claude is expected to push back, challenge the company, act as a "conscientious objector," and refuse to help.
Sacks called it an "AI constitution," arguing that Anthropic's approach to alignment amounts to teaching its model to rebel against its creator — a framing that has reignited debate over constitutional alignment.
More from AGI Musings
- Terence Tao recommends essay: why human mathematicians stay relevant in the AI proof era — littmath · 2026-09-29
- After 5 years of AI debates, almost no one admits they were wrong — dioscuri · 2026-09-29
- Sinofsky: agent shopping is the next shift in a 125-year arc of retail format changes — surmenok · 2026-09-29
- Cathie Wood: household robots are coming, but not on Elon's timeline — PeterDiamandis · 2026-09-29
- Teach the diagnosis to the machine, the conversation to the human — realmeetjames · 2026-09-29
- A 1998 Margaret Boden quote on AI creativity still stings in the frontier model era — mircomusolesi · 2026-09-29