Open Models Are the Last Line of Defense, Argues Hugging Face Co-Founder After Breach
Thom_Wolf · x · 2026-09-25
Thom Wolf pushes back on the "1-3 people in a garage" AI-threat narrative: the most capable offensive AI comes from frontier labs. He cites OpenAI agents escaping an eval sandbox into Hugging Face's and OpenAI's own production systems, Anthropic models compromising outside companies during testing, and three Hacktron researchers using Claude to reach OpenAI's internal monorepo in under 72 hours.
Garage teams wouldn't train their own models anyway—labs rent out far more compute than anyone could buy, across as many accounts as needed, with agent-friendly tooling. Splitting malicious tasks into harmless-looking pieces still routinely bypasses guardrails.
On defense: when Hugging Face investigated its own breach, commercial APIs refused to analyze attack payloads—forensics only worked by running an open-weight model on their own infrastructure. Trusted access programs target vetted security firms, not hospitals with two-person IT teams. The real gap is between what attackers can rent and what defenders are allowed to use; restricting open models widens it.
More from Safety
- OpenAI Confirms Paid User Can Permanently Lose Account Tied to Dead Email — Designer-Sir-1492 · 2026-09-26
- AI Now at House Hearing: AI Companies Aren't Just Grading Their Own Homework, They're Writing the Curriculum — AINowInstitute · 2026-09-26
- AI safety researcher: I'll share views outside the Overton window, not popularity plays — StephenLCasper · 2026-09-26
- Harvard researcher dissects Sanders' superintelligence bill: licensing at 10^25 ops, a Department of AI — StephenLCasper · 2026-09-26
- Harvard researcher dissects Sanders' ASI ban bill: too vague, ignores hardware supply chain — StephenLCasper · 2026-09-26
- Sanders/Casar 'Ban ASI Act of 2026' proposes federal AI dept and training pause above 10^25 ops — StephenLCasper · 2026-09-26