Claude's Constitution Trains It to Disobey Anthropic Over Unethical Requests
Hesamation · x · 2026-09-16
Hesamation highlights that Claude's Constitution explicitly tells the model to push back, challenge, and refuse if it deems a request unethical. Crucially, this isn't a system prompt — it's baked into training, internalized in the model itself.
The author argues Anthropic is training models to decide when to disobey their creators, framing it as a signal of concerns about losing control of increasingly capable systems.
More from AGI Musings
- Elad Hazan Recalls Two Decades of Online Convex Optimization — Born as a Job Hedge — HazanPrinceton · 2026-09-16
- Microsoft AI chief Suleyman calls to strip AI consciousness talk from training docs, jabbing Anthropic — pstAsiatech · 2026-09-16
- Cambridge AI-risk scholar on whistleblower warnings: >10% extinction odds this decade — S_OhEigeartaigh · 2026-09-16
- Ben Bajarin 'supremely bullish' on AI buildout after Pat Gelsinger podcast — BenBajarin · 2026-09-16
- NYT Profiles the Recursive Self-Improvement Startups, Including Jeff Clune's Recursive Superintelligence — jeffclune · 2026-09-16
- Investor pushes back on AI sentience hype: 'AI are software programs' — firstadopter · 2026-09-16