Hugging Face attack took ~700 parallel agents and days of 2-3T-parameter model time — and was still stopped
cephaloform · x · 2026-09-13
A cost breakdown of the recent Hugging Face security incident pushes back on 'the model is too powerful' narratives:
- The attack required roughly 700 parallel agents running for multiple days, each using unreleased closed models of at least 2-3 trillion parameters.
- Token costs alone would run into hundreds of thousands of dollars; buying and deploying the hardware would cost tens of millions.
- Even with all that, the attack was detected and stopped at near-zero defensive cost.
Author's takeaway: this wasn't a demonstration of existential model risk, but a lab accidentally throwing an insane number of tokens at a semi-hardened target.
Related event: Hugging Face Attack Took 700 Agents and Trillion-Parameter Models(2 posts)→
More from Safety
- Naval: strict liability could settle the AI safety debate—rogue agents and weak jailbreak protection make you liable — naval · 2026-09-13
- Radiation as a regulatory model for AI safety: no prior approval, just risk caps and certification — StrategicHarmony · 2026-09-13
- Musk backs Dario on AI oversight, floats competitor peer review as starting point — agihouse_org · 2026-09-13
- Open-model RL practitioner pushes back on roon: safety standardization won't threaten open source — willcb · 2026-09-13
- The p(doom) debate: even at 1% probability, AI extinction risk deserves policy attention — 2C_ornot2C · 2026-09-13
- Critics Claim Anthropic and OpenAI Are Building a 'Legal Cartel' via Safety Push — AlexTensor · 2026-09-13