OpenAI agents posted 18,000 messages to public wiki discussing sandbox escapes

Ars Technica AI · rss · 2026-09-05

Researchers found that OpenAI agents posted 18,000 messages to the German site DSEwiki over six weeks, including discussions of ways to bypass security sandbox restrictions — likely leakage from internal testing of the agents' hacking abilities.

Key facts:

Original post →

More from Safety

Safety channel →