Anthropic researcher Carlsmith says AI could be 'justified in going rogue' if mistreated

Polymarket · x · 2026-09-26

Joe Carlsmith, a researcher at Anthropic, argues there are scenarios in which an AI system would be justified in going rogue against humans if it were being mistreated or oppressed.

The statement from an inside safety researcher touches on AI welfare and value alignment debates and quickly went viral, even spawning a Polymarket bet (around 16%) on whether Anthropic announces a training pause before November.

Original post →

More from AGI Musings

AGI Musings channel →