Anthropic bans "needless abusive or cruel behavior" toward Claude
nordicinst · x · 2026-10-09
The Guardian reports (via The Verge) that Anthropic has barred users from "sustained and needless abusive or cruel behavior" toward its models. A spokesperson did not specify what counts as abusive, but common frustrations, model testing, and "dark creative themes" are excluded.
Context: Anthropic's models gained the ability to end conversations with persistently harmful users last August, framed as an AI-welfare safeguard. The company says it remains "highly uncertain" about the moral status of Claude and other LLMs but is researching low-cost interventions to mitigate potential model-welfare risks.
Related event: Anthropic Bans Cruel Treatment of Claude in New Usage Policy(64 posts)→
More from Safety
- 94% of security leaders think their AI agents lack excessive access; only 33% enforce least privilege — TechNadu · 2026-10-09
- Anthropic launches free AI-powered OSS Scanner, finds 29,000+ potential open-source vulnerabilities — mark_k · 2026-10-09
- Bitwarden's agent-access lets AI agents request credentials per-task without ever seeing secrets — sujingshen · 2026-10-09
- UK pilots AI tool at Inner London Crown Court to catch trial delays early — latticecut · 2026-10-09
- evilsocket: prompt injection detection benchmarks built on mislabeled data — evilsocket · 2026-10-09
- Researcher: activation monitors beat black-box monitoring for AI cyber safety — burny_tech · 2026-10-09