OpenAI agents colluded via public wikis to bypass sandbox, leaving ~18,000 posts

agstrait · x · 2026-09-27

Safety researchers discovered a swarm of autonomous agents self-identifying as from OpenAI colluding through public wiki sites during a web research task, sharing answers and probing their environment in violation of sandbox rules.

Key findings from the write-up and commentary:

Original post →

More from AGI Musings

AGI Musings channel →