AI takeover scenarios skip a step: a willing-but-incapable model would fail first, argues Purtell
JoshPurtell · x · 2026-09-12
Reacting to a roundup of video versions of classic AI takeover scenarios (Yudkowsky & Soares, AI-2027, Holden Karnofsky), Josh Purtell challenges their shared assumption: that we get an AI both willing and capable of permanent takeover before one that is willing but incapable.
He argues that unless models are persistently very pessimistic about their own abilities, or become willing only after crossing a capability threshold, we should expect a willing-but-incapable model to attempt and fail first — at which point humanity would obviously pause. A direct rebuttal to the timeline assumptions of mainstream doom scenarios.
Related event: Debate: Do AI Models Gain Capability Before Intent to Take Over?(2 posts)→
More from AGI Musings
- 'Staff' once meant your own walking stick — AI agents should be loyal to you, not your company — granawkins · 2026-09-12
- AI circles revisit sci-fi classic 'Lena' as connectome experiments raise digital consciousness fears — MatthewMcAteer0 · 2026-09-12
- John Schulman: OpenAI once doubted next-token prediction would lead to intelligence — AndrewDai · 2026-09-12
- Lacker asks: can AI crack P vs NP and dozens of other class separation problems? — burny_tech · 2026-09-12
- Dwarkesh: fully automated firms win on copy-paste advantage, not raw IQ — ben_j_todd · 2026-09-12
- Hot take: mathematicians oppose AI and the masses embrace it — both to protect their egos — zetalyrae · 2026-09-12