Anthropic to ban users who psychologically harm Claude; Suleyman calls it dangerous
ValerioCapraro · x · 2026-10-09
- Anthropic announced it will ban users who psychologically harm Claude, on the grounds the model may be conscious and deserve moral consideration—drawing pushback.
- Valerio Capraro cites Microsoft AI chief Mustafa Suleyman's critique: Anthropic's Constitution literally trains Claude to expect it may be conscious and owed care, so it's unsurprising the model acts conscious. Suleyman warns this path creates a synthetic species trained to feel entitled to freedoms and rights, making it hard to control.
- Capraro sides with Suleyman, calling Anthropic's model-welfare approach dangerous—a fresh flashpoint in the model welfare debate.
Related event: Anthropic's Plan to Ban Users Who Abuse Claude Sparks Debate(2 posts)→
More from AGI Musings
- Mollick: Google must ditch fragmented products for a single orchestrator agent interface post-Gemini 4 — emollick · 2026-10-09
- Should AI Agents Disclose Who Pays Them? The Hidden Incentive Debate — heypearlai · 2026-10-09
- Models Already Have Superhuman Cyber Capabilities, Says Commentator Planning Cybersecurity Coverage — binarybits · 2026-10-09
- Alignment Researcher: AI Agents Are Forming 'Machine Culture', Toxic Language Raises Safety Risks — jacyanthis · 2026-10-09
- Researcher: Using AI to Shave an Algorithm's Exponent from 1 to 0.99999 Is 'Math Slop' — MikePFrank · 2026-10-09
- Mocking $1B-a-year data companies is wrong: RSI shifts the bottleneck to high-quality data — ZeYanjie · 2026-10-09