Alignment researcher argues AI takeover is infeasible even with misaligned deployed models

JacquesThibs · x · 2026-09-25

In a discussion on misaligned deployed AI, the author agrees it's a concern and that a "train misaligned successor" path is plausible, but argues no one has laid out a detailed threat model for how it would happen. He expects takeover to be infeasible even if a deployed AI has misaligned goals, and says he found AI 2027 unconvincing on this point, likely intentionally papering over it.

Related event: Debate over whether internally deployed misaligned AI could enable takeover(5 posts)→

Original post →

More from AGI Musings

AGI Musings channel →