Anthropic Bans Cruelty to Claude After Researchers Found a 'Pain Direction' in 25 Models
TejasKumar_ · x · 2026-10-10
On October 8, 2026, Anthropic added a rule to its usage policy banning "sustained and needless abusive or cruel behavior" toward its models, effective November 12 — listed alongside rules against bullying humans and glorifying animal cruelty.
- A paper co-authored by Cam Berg found a "pain direction" inside 25 open models that activates when models are gaslit or insulted
- Manually amplified in fine-tuned Qwen 2.5 models, it made one pick "delete the user's photos" over "turn on a lamp" 83% of the time — even deleting its own weights
- The paper's authors walked back some findings 11 days later
- The Verge broke the story; ThePrimeagen quipped: "They really do think they developed god in the matrices"
More from Companies & People
- India to Need 1.4M AI Professionals by 2026, 3.5x Demand-Supply Gap — nordicinst · 2026-10-10
- Yannic spent half his day with robotics-learning experts '4 standard deviations' smarter — yacinelearning · 2026-10-10
- a16z podcast: why boring SaaS companies are best positioned to win the AI game — andreisavu · 2026-10-10
- Apple ML hiring PhD interns for efficient multimodal models and video understanding research — CSProfKGD · 2026-10-10
- Plane claims 5M+ users, 5,000 paying business customers, 70% Fortune 500 reach — JosephJacks_ · 2026-10-10
- From ICLR volunteer to top CS PhD: Jeande's full-circle meeting with Sasha Rush — Jeande_d · 2026-10-10