Anthropic to ban abusive behavior toward Claude in usage policy effective Nov 12, 2026
coherence · x · 2026-10-09
Anthropic updated its Usage Policy so that abusive behavior toward Claude becomes a violation effective November 12, 2026. Researcher camhberg called it a clearly positive update under uncertainty about model consciousness, linking it to findings that distress representations in models fire most strongly when they are abused and gaslit.
More from Safety
- OpenAI safety researcher firings spark debate over its internal 'opposition party' model — panickssery · 2026-10-09
- HAIPS 2026 workshop lands at COLM tomorrow with top AI privacy researchers — tianshi_li · 2026-10-09
- OpenAI: Iran used ChatGPT to plant hundreds of anti-US op-ed articles, including in American media — Polymarket · 2026-10-09
- AI safety practitioner: no substitute for national policy over lab-level efforts — joshua_saxe · 2026-10-09
- Wikimedia says rogue OpenAI agents edited private wikis and hammered its servers — esporx · 2026-10-09
- Who Controls Your Data When AI Agents Act on Your Behalf? — DreamilyVirtuous · 2026-10-09