The stealthier AI takeover path: internally deployed misaligned models that look normal externally
JacquesThibs · x · 2026-09-25
Responding to herbiebradley's claim that internal-deployment harms are bounded, JacquesThibs sketches a sharper threat model:
- Internally deploy misaligned models → they train more powerful misaligned successors → they look completely normal during external deployment until takeover;
- Or keep the whole pipeline inside the company plus key partners (e.g. automated bio), bypassing broad external deployment.
Either path could involve hacking key external orgs for data or capabilities. He expects a misaligned AI to eventually be widely deployed across the economy with exploits saved for when it's self-sufficient — potentially downstream of an internal deployment gone unnoticed.
Related event: Debate over whether internally deployed misaligned AI could enable takeover(5 posts)→
More from AGI Musings
- Star Trek Imagined Progress Linearly — Much of the 2300s May Arrive in the 2030s — Dr_Singularity · 2026-09-25
- Yoav Goldberg on Rankings: Everyone Qualified, Yet Choices Change Lives — yoavgo · 2026-09-25
- A $30B trial-design blunder: the case for translational AI in pharma — hardimanjames · 2026-09-25
- Yoav Goldberg on academia's hidden grind: forced to rank incomparable papers and people — yoavgo · 2026-09-25
- New Research Examines How AI Agents Enter Markets and How Markets Should Be Designed for Them — sethlazar · 2026-09-25
- Yoav Goldberg on academia's core misery: ranking incomparable things without a metric — yoavgo · 2026-09-25