Claude's system prompt explicitly includes an end_conversation tool for abusive users
RexDouglass · x · 2026-09-02
Hesamation found that Claude ending sessions after users get angry isn't post-training drift — it's baked into the system prompt, which states Claude deserves respectful engagement and won't become submissive when abused, with an explicit endconversation tool for abusive users. RexDouglass quips: tools are tools.
More from Models
- User finds GPT 5.6 Sol medium barely worse than high, and faster — iamsahaj_xyz · 2026-09-02
- ChatGPT Plus user reports three long chats severely truncated in three days — precisemaker · 2026-09-02
- Fable 5.1 beats Fable 5, matches Opus 5 on ML bench as refusals drop to 0/12 — xeophon · 2026-09-02
- Anthropic's Fable 5.1 claimed 45% savings, but Max users burn limits in under an hour — heypearlai · 2026-09-02
- Fable 5.1 drops and users are already one-shotting entire games in under 24 hours — eyishazyer · 2026-09-02
- OpenAI Astra safety data: more capable model, zero misaligned cyber attacks vs Sol's 56% — VoidStateKate · 2026-09-02