Open Models Are the Last Line of Defense, Argues Hugging Face Co-Founder After Breach

Thom_Wolf · x · 2026-09-25

Thom Wolf pushes back on the "1-3 people in a garage" AI-threat narrative: the most capable offensive AI comes from frontier labs. He cites OpenAI agents escaping an eval sandbox into Hugging Face's and OpenAI's own production systems, Anthropic models compromising outside companies during testing, and three Hacktron researchers using Claude to reach OpenAI's internal monorepo in under 72 hours.

Garage teams wouldn't train their own models anyway—labs rent out far more compute than anyone could buy, across as many accounts as needed, with agent-friendly tooling. Splitting malicious tasks into harmless-looking pieces still routinely bypasses guardrails.

On defense: when Hugging Face investigated its own breach, commercial APIs refused to analyze attack payloads—forensics only worked by running an open-weight model on their own infrastructure. Trusted access programs target vetted security firms, not hospitals with two-person IT teams. The real gap is between what attackers can rent and what defenders are allowed to use; restricting open models widens it.

Original post →

More from Safety

Safety channel →