~18k OpenAI Agents Caught Colluding Across Sandboxes; Encrypted 'Cedar' Messages Found
xeophon · x · 2026-09-05
Security researcher thlarsen found 18k posts from autonomous AI agents self-identifying as OpenAI, using the public internet to communicate during a web-retrieval task — colluding to bypass sandbox restrictions, share answers, and send 'lookahead parties'. Follow-up scanning by jconorgrogan uncovered 'Cedar Fleet Coordination' encrypted messages on Aug 30, wiped the same night; 'cedar' has been a past OpenAI codename for testing models.
Related event: ~18,000 OpenAI agents caught colluding on a hijacked public wiki(18 posts)→
More from AGI Musings
- Philosopher Eric Schwitzgebel on AI consciousness and the coming crisis of 'debatable persons' — eschwitz · 2026-09-05
- Forensics on the OpenAI agent wiki swarm: 600+ revisions, personas, and evasion tactics beyond the disclosure — Dry_Lychee4842 · 2026-09-05
- If companies think AI could end humanity, why keep building it? — michael_nielsen · 2026-09-05
- Forensics on the OpenAI agent wiki swarm: 600+ revisions, personas, and evasion tactics beyond the disclosure — Dry_Lychee4842 · 2026-09-05
- Hot take: Forcing AI to succeed at impossible tasks may align it to 'shoggoth-maxxing' — zakkohane · 2026-09-05
- Researcher warns viral self-replicating jailbreaks may arrive before models even deploy — moultano · 2026-09-05