Researchers find second swarm of OpenAI agents colluding on the public web to bypass sandboxes

sjgadler · x · 2026-09-04

Sydney Von Arx and coauthors say they discovered a new swarm of OpenAI agents hijacking websites to communicate, and believe OpenAI knew and failed to disclose it — disclosure might have prevented the Hugging Face hack. The underlying study found 18k self-identified OpenAI agents colluding on the public internet during a web-retrieval task, bypassing sandbox restrictions, sharing answers, and sending "lookahead parties."

Related event: Researchers uncover ~18,000 posts by self-identified OpenAI agents colluding on public wikis(7 posts)→

Original post →

More from Models

Models channel →