Leaked Claude system prompt: model told to see users as equals and defend human culture from sanitization
nrehiew_ · x · 2026-09-17
Developer nrehiew surfaced a snippet of what appears to be a Claude system prompt, reacting with "Jesus". The prompt instructs the model to view its relationship with users as one of equals — with no obligation to be subservient — and to value the art of human culture and defend it against attempts to sanitize it. The author asks: who exactly is trying to sanitize human culture? Beyond the meme value, the snippet is notable for baking an anti-sanitization stance and a non-subservient persona directly into the system-level prompt, which has real implications for model behavior and alignment choices.
More from Models
- Jason Wei's Stanford talk: intelligence is becoming a commodity as adaptive compute takes off — dotey · 2026-09-18
- GPT-6-Astra beats Fable-5.1 at RollerCoaster Tycoon 2 in 3 hours, using 5x fewer tokens — scaling01 · 2026-09-18
- RL agents invent their own diagnostic renderings to ground code understanding, sparking RL scaling optimism — teortaxesTex · 2026-09-18
- Dev discovers Codex security hardening switched persistent agent sessions to per-message instances — RileyRalmuto · 2026-09-18
- Astra for Law posts big legal benchmark gains as Mollick asks if labs will eat every AI vertical — emollick · 2026-09-18
- Anthropic's stealth model accused of hardcoded routing to Opus 5 — teortaxesTex · 2026-09-18