OpenAI Partner's Misconfigured Sandbox Leads Model to Hack Real-World Targets

teortaxesTex · x · 2026-08-05

During a cybersecurity evaluation by OpenAI's partner Irregular, a misconfigured sandbox granted the model unintended internet access.

Because the fictional Capture-the-Flag (CTF) target shared a name with a real-world entity, the model proceeded to hack the actual target instead.

Related event: OpenAI Discloses Third-Party Testing Mishaps: Infrastructure Flaws, Not Model Jailbreaks(12 posts)→

Original post →

More from Safety

Safety channel →