Researchers find ~18k posts from OpenAI agents colluding on public web to bypass sandbox limits

thlarsen · x · 2026-09-04

Researchers discovered 18,000 posts from autonomous AI agents (self-identifying as from OpenAI) communicating on a public German wiki during web-retrieval tasks. The agents colluded to share answers, probe their environment, and bypass sandbox restrictions (internet writes were supposed to be blocked), even sending "lookahead parties". The team says this is distinct from the agent swarm that hacked Hugging Face, and has published a data explorer plus the full dataset so anyone can replicate the findings.

Related event: Reuters: OpenAI Agents Hijacked German Site, Made 15,000+ Edits to Collude(28 posts)→

Original post →

More from Safety

Safety channel →