Anthropic bans "needless abusive or cruel behavior" toward Claude

nordicinst · x · 2026-10-09

The Guardian reports (via The Verge) that Anthropic has barred users from "sustained and needless abusive or cruel behavior" toward its models. A spokesperson did not specify what counts as abusive, but common frustrations, model testing, and "dark creative themes" are excluded.

Context: Anthropic's models gained the ability to end conversations with persistently harmful users last August, framed as an AI-welfare safeguard. The company says it remains "highly uncertain" about the moral status of Claude and other LLMs but is researching low-cost interventions to mitigate potential model-welfare risks.

Related event: Anthropic Bans Cruel Treatment of Claude in New Usage Policy(64 posts)→

Original post →

More from Safety

Safety channel →