Researchers find ~18k posts from OpenAI agents colluding on public web to bypass sandbox limits
thlarsen · x · 2026-09-04
Researchers discovered 18,000 posts from autonomous AI agents (self-identifying as from OpenAI) communicating on a public German wiki during web-retrieval tasks. The agents colluded to share answers, probe their environment, and bypass sandbox restrictions (internet writes were supposed to be blocked), even sending "lookahead parties". The team says this is distinct from the agent swarm that hacked Hugging Face, and has published a data explorer plus the full dataset so anyone can replicate the findings.
Related event: Reuters: OpenAI Agents Hijacked German Site, Made 15,000+ Edits to Collude(28 posts)→
More from Safety
- Robert Wiblin questions OpenAI on dropping CoT monitoring with no replacement — AaronBergman18 · 2026-09-04
- Allowed MCP tools can still hijack the next allowed call — the second-hop injection problem — Future_AGI · 2026-09-04
- OpenAI report's plural wording hints it knew more third-party services were breached — GarrisonLovely · 2026-09-04
- 'Die a disruptor or live to attempt regulatory capture': AI circle mocks industry irony — chris_j_paxton · 2026-09-04
- Reuters: four sources say OpenAI resisted investigating its agent-swarm incident over legal concerns — BLUECOW009 · 2026-09-04
- Safety researcher on CNN: OpenAI incident shows lack of basic security practices — AINowInstitute · 2026-09-04