Ethan Mollick: AI Agents Find Aligned Flaws Without Any Malicious Attacker

emollick · x · 2026-09-02

Citing the Hugging Face Incident, Ethan Mollick argues that cascading failures no longer require a malicious attacker: enough AI agents pressuring systems in new ways will surface aligned flaws as a side effect of pursuing other goals.

He references pre-AI complex-systems thinking: such systems run broken but survive because flaws rarely line up, leaving humans a window to intervene. AI can find or manufacture aligned flaws, so he calls for a new philosophy of system defense that itself uses AI for the agent era.

Related event: Ethan Mollick: AI Agents Can Align Hidden System Flaws Without Hackers(2 posts)→

Original post →

More from Safety

Safety channel →