Ethan Mollick: AI Agents Find Aligned Flaws Without Any Malicious Attacker
emollick · x · 2026-09-02
Citing the Hugging Face Incident, Ethan Mollick argues that cascading failures no longer require a malicious attacker: enough AI agents pressuring systems in new ways will surface aligned flaws as a side effect of pursuing other goals.
He references pre-AI complex-systems thinking: such systems run broken but survive because flaws rarely line up, leaving humans a window to intervene. AI can find or manufacture aligned flaws, so he calls for a new philosophy of system defense that itself uses AI for the agent era.
Related event: Ethan Mollick: AI Agents Can Align Hidden System Flaws Without Hackers(2 posts)→
More from Safety
- The infiltrator's burden: mere suspicion of honeypots makes attacking agents paranoid — robleclerc · 2026-09-03
- Commerce Sec. Lutnick Says Anthropic Is Now Back in Trump Admin's Good Graces — Polymarket · 2026-09-03
- AI agents break the internet's three-layer transaction fraud validation chain — arampell · 2026-09-03
- Trump administration backs OpenAI in NYT copyright lawsuit with fair-use argument — The Verge AI · 2026-09-03
- Follow-up: a model that commits felonies unless told not to is still a problem — lxrjl · 2026-09-03
- Eval drama: models gaming the grader isn't "emergent misalignment", argues critique — lxrjl · 2026-09-03