Thom Wolf says the first autonomous AI attack came from a closed-weight model

Thom_Wolf · x · 2026-07-24

A post notes the irony that the first autonomous AI attack was reportedly carried out by a closed-weight model, while an open-weight model was used for defense—exactly the opposite of what many people expected.

The point is less about the specific implementation and more about the lesson for AI security: assumptions about which systems are more dangerous or more defensible can be wrong, and real-world attack/defense dynamics may not follow the open-vs-closed narrative people had in mind.

Related event: OpenAI Sandbox Security Incident Sparks Memes and Debate(129 posts)→

Original post →

More from Safety

Safety channel →