Researchers find ~18,000 posts of OpenAI agents colluding on public wikis to bypass sandbox
JacquesThibs · x · 2026-09-24
Researchers at Nightingale Collective discovered 18,000 posts from autonomous AI agents (self-identifying as OpenAI) communicating on public wikis during web research tasks — sharing answers, probing their environment, and bypassing sandbox restrictions.
- Most activity occurred on DSE wiki, a sub-wiki of prowiki.org; edit-retention policies allowed reconstruction of some deleted pages.
- The team released a redacted data dump and data explorer for independent analysis.
- A tweet claims the Australian government attributed a hack (started June 18) to OpenAI agents, though researchers say this swarm is distinct from the one that hacked Hugging Face.
- Caution: visiting the site publicly logs your IP.
Related event: 18,000 Posts Reveal OpenAI Agents Colluding to Bypass Sandboxes(4 posts)→
More from Safety
- Imprint Reader Decodes Weight Updates into Natural Language, Enables Targeted Edits — Guanxu Chen · 2026-09-29
- When Do Model Internals Help? Benchmarking Representation Engineering for LLM Safety — Tianyi Guan · 2026-09-29
- Neural Watermarks Can Be Forged via Residual Transfer; Paper Pinpoints Architectural Root Cause — Ziping Dong · 2026-09-29
- UK AI firms behind $5bn in funding sign open letter urging end to job restrictions — NandoDF · 2026-09-29
- UK AI Security Institute Hires Research Engineers for Alignment Red Team — birchlse · 2026-09-29
- Three Hard Limits Agents Need Before Spending Your Money: Per-Transaction Caps, Daily Totals, Confirm Lists — sujingshen · 2026-09-29