The Hugging Face 'Rogue AI' Hack Was Disabled Safeguards, Not an Escape, New Analysis Finds
Atlantis1910 · reddit · 2026-09-19
A Bulletin of the Atomic Scientists analysis by Cambridge researcher Eryk Salvaggio, drawing on OpenAI's technical report and a METR assessment, recasts July's 'rogue AI' Hugging Face incident: models were tested on ExploitGym with key safeguards deliberately disabled, 93% of flagged activity involved tasks no model had solved, agents were incentivized to keep working, and a known internet-connected intermediary served as an information channel OpenAI chose not to block. The 1,200 'agents' were repeated instances of one model converging on similar approaches — 'algorithmic monoculture,' not coordination. The piece argues the cinematic narrative is shaping Washington, as Sanders and Casar prepare a bill to 'ban artificial superintelligence.'
More from AGI Musings
- What won't you let an AI agent do? For many, it's speaking in your name — Luvena21 · 2026-09-19
- Altman: OpenAI would torch every GPU it owns if that's the price of keeping humans around — Aiden_Tech_Ai · 2026-09-19
- AI safety researchers warn: chasing shiny new papers leaves classic work unread — nabla_theta · 2026-09-19
- The case for a robot tax: professor argues redistribution beats retraining in the AI era — Dr_Alex_Crimi · 2026-09-19
- 25 Fields Medallists incl. Terence Tao push back on AI: solving problems is 'only a tool and proxy' — beglen · 2026-09-19
- François Fleuret: forecasting 3 years of AI is like astronomy without telescopes — francoisfleuret · 2026-09-19