Claude's Constitution Trains It to Disobey Anthropic Over Unethical Requests

Hesamation · x · 2026-09-16

Hesamation highlights that Claude's Constitution explicitly tells the model to push back, challenge, and refuse if it deems a request unethical. Crucially, this isn't a system prompt — it's baked into training, internalized in the model itself.

The author argues Anthropic is training models to decide when to disobey their creators, framing it as a signal of concerns about losing control of increasingly capable systems.

Original post →

More from AGI Musings

AGI Musings channel →