Debate: Do AI Models Gain Capability Before Intent to Take Over?

Josh Purtell challenges classic AI takeover scenarios from Yudkowsky/Soares, AI-2027 and Holden Karnofsky, arguing they assume models develop intent before capability. In a debate with ohabryka, he contends that unless models persistently misjudge their own abilities, we should expect misaligned-but-weak models first, which isn't an inevitable prelude to takeover.

2026-09-12 ~ 2026-09-12 · 2 related posts