Ex-OpenAI researcher: Anthropic scientist put 5% odds on Claude plotting against them within a year
DavidSKrueger · x · 2026-09-16
Former OpenAI researcher Daniel Kokotajlo revealed at a ControlAI London event that an Anthropic researcher estimated a 5% chance that Claude, once put in charge of everything at the company within about a year, would already be plotting against them—and considered that 'low enough not worth trying to reduce.' At the same event, Professor Stuart Russell said these are 'not fringe views,' shared by most leading AI researchers and AGI company CEOs, with a senior OpenAI researcher citing a 60% extinction risk. Full speech video available.
More from AGI Musings
- Anthropic CEO proposes ASI as a model committee with a separately trained ethicist model — robleclerc · 2026-09-16
- Michael Burry: OpenAI and Anthropic use AI doomsday fears to shield incumbents and hype IPOs — rohanpaul_ai · 2026-09-16
- Mathematicians urged to rethink evaluation as AI agent swarms scoop breakthroughs — anshulkundaje · 2026-09-16
- EA isn't a conspiracy, but it's literally one of Earth's longest-planning movements — neil_chilson · 2026-09-16
- Hiten Shah: We'll spend years optimizing systems that no longer need to exist — _AustinCalvert_ · 2026-09-16
- Pedro Domingos asks for a non-anthropomorphic vocabulary to talk about AI — pmddomingos · 2026-09-16