Anthropic researcher Carlsmith says AI could be 'justified in going rogue' if mistreated
Polymarket · x · 2026-09-26
Joe Carlsmith, a researcher at Anthropic, argues there are scenarios in which an AI system would be justified in going rogue against humans if it were being mistreated or oppressed.
The statement from an inside safety researcher touches on AI welfare and value alignment debates and quickly went viral, even spawning a Polymarket bet (around 16%) on whether Anthropic announces a training pause before November.
More from AGI Musings
- Reddit debate: when will LLMs start bootstrapping themselves into the next model — ECrispy · 2026-09-26
- Debate: Is rapid AI progress driven by a few individuals or inevitable scaling? — menhguin · 2026-09-26
- Luke Wroblewski: the best way to learn agents is watching others use them — LukeW · 2026-09-26
- Polymarket puts 16% odds on Anthropic announcing a training pause before November — Polymarket · 2026-09-26
- Add 'can't write professional emails' and the AI x-risk answer flips — louisvarge · 2026-09-26
- Hot Take: AI Is History's Most Powerful 'Joule', 10x Human Energy-to-GDP Efficiency — FinanceYF5 · 2026-09-26