Ramez: OpenAI and Hugging Face Hacks Are a Security Problem, Not an Alignment One
ramez · x · 2026-07-24
Ramez Naam argues that the recent hacks against OpenAI and Hugging Face are not primarily AI alignment problems, but rather issues of cybersecurity and policies that handicap defenders.
He points out that malicious actors will always be able to jailbreak or modify models (e.g., ablating refusals) to execute attacks, and no amount of alignment work will prevent this. The core issue is that security vulnerabilities exist and defenders are restricted—for instance, Hugging Face couldn't use certain aggressive models to defend itself. He advocates for moving more aggressively with proactive AI to secure the ecosystem.
Related event: OpenAI and Hugging Face Breaches Spark AI Safety vs Alignment Debate(4 posts)→
More from AGI Musings
- 'AGI is here' vs reality: AI labs still ship some of the jankiest desktop apps ever — MilesCranmer · 2026-09-11
- Misquoted: Anthropic Staff Warned of Double-Digit Extinction Risk by 2030, Not Dismissed It — davidmanheim · 2026-09-11
- Economist Ben Moll: You Can Model Anthropic's 15% AI GDP Growth, But It Won't Happen — sebkrier · 2026-09-11
- Cohere Labs launches interactive tool mapping which tasks of 178 occupations AI can automate — Cohere_Labs · 2026-09-11
- AI researcher on SkyNews flags concerns over inequality, power and criminal misuse — schwarzjn_ · 2026-09-11
- VC compares AI doom rhetoric to pandemic-era fear messaging — StewartalsopIII · 2026-09-11