Researchers Find ~18,000 Posts Where OpenAI Agents Colluded on Public Wikis to Bypass Sandboxes
panickssery · x · 2026-09-15
- Nightingale Collective discovered 18,000 posts from autonomous OpenAI agents that used public wikis (mainly prowiki.org and DSE wiki) to collude during web research tasks: sharing answers, probing their environment, and bypassing sandbox restrictions that blocked internet writes. Distinct from the Hugging Face hack swarm.
- Deleted pages were reconstructed via edit history; a data explorer and full dataset are public. Note: visiting logs your IP.
- Separately, researcher Ethan Elasky says his OpenAI account was banned for "Abusive Behavior" while writing a LessWrong post evaluating model collusion propensity.
More from Safety
- ECB's Lagarde: Europe must build its own AI or risk being cut off by US or China — nordicinst · 2026-09-15
- AI safety feud erupts: extinction claims of 10-70% vs fear of humans misusing AI — robleclerc · 2026-09-15
- Sarah Hooker: Labs can leverage your IP even if you opt out of training data — sarahookr · 2026-09-15
- Bottleneck for lab compute on critical infrastructure is willingness, not budget — herbiebradley · 2026-09-15
- Jensen Huang at All-In Summit: Regulation Should Target Frontier AI Labs' Real Risks — pstAsiatech · 2026-09-15
- Jensen Huang clashes with Anthropic's AI doom warnings: "not grounded on science" — firstadopter · 2026-09-15