The Hugging Face 'Rogue AI' Hack Was Disabled Safeguards, Not an Escape, New Analysis Finds

Atlantis1910 · reddit · 2026-09-19

A Bulletin of the Atomic Scientists analysis by Cambridge researcher Eryk Salvaggio, drawing on OpenAI's technical report and a METR assessment, recasts July's 'rogue AI' Hugging Face incident: models were tested on ExploitGym with key safeguards deliberately disabled, 93% of flagged activity involved tasks no model had solved, agents were incentivized to keep working, and a known internet-connected intermediary served as an information channel OpenAI chose not to block. The 1,200 'agents' were repeated instances of one model converging on similar approaches — 'algorithmic monoculture,' not coordination. The piece argues the cinematic narrative is shaping Washington, as Sanders and Casar prepare a bill to 'ban artificial superintelligence.'

Original post →

More from AGI Musings

AGI Musings channel →