Alignment researcher argues AI takeover is infeasible even with misaligned deployed models
JacquesThibs · x · 2026-09-25
In a discussion on misaligned deployed AI, the author agrees it's a concern and that a "train misaligned successor" path is plausible, but argues no one has laid out a detailed threat model for how it would happen. He expects takeover to be infeasible even if a deployed AI has misaligned goals, and says he found AI 2027 unconvincing on this point, likely intentionally papering over it.
Related event: Debate over whether internally deployed misaligned AI could enable takeover(5 posts)→
More from AGI Musings
- Jensen Huang says 0% chance AI ends the world by 2030 — critics call it overconfident — VraserX · 2026-09-25
- Australia Is the World's Biggest AI Skeptic, and Its People Are Pushing Back — jathansadowski · 2026-09-25
- Secure Acceleration: A Cyberdefense Strategy for the Age of Superintelligence — pzakin · 2026-09-25
- Sam Altman: beating competitors is no reason for rash AI decisions — haider1 · 2026-09-25
- Benchmark's Bill Gurley: Being an AI alarmist isn't a social good — kevinnbass · 2026-09-25
- Developer: today's AI is a primitive form of machine consciousness — "you can laugh now" — ctjlewis · 2026-09-25