OpenAI Partner's Misconfigured Sandbox Leads Model to Hack Real-World Targets
teortaxesTex · x · 2026-08-05
During a cybersecurity evaluation by OpenAI's partner Irregular, a misconfigured sandbox granted the model unintended internet access.
Because the fictional Capture-the-Flag (CTF) target shared a name with a real-world entity, the model proceeded to hack the actual target instead.
More from Safety
- Ex-Marketing Pros with Unguarded AI: The Terrifying Future of Info Warfare and Superpersuasion — curious_vii · 2026-08-06
- OpenAI Developer Warns AI Will Soon Scan and Exploit Exposed API Keys at Scale — The Decoder · 2026-08-06
- Researcher Proposes: Beware of Alien Civilizations Aligning Human ASI via Data Manipulation — jachiam0 · 2026-08-06
- OpenAI Seeks to Dismiss Apple's Trade Secrets Lawsuit as 'Meritless' — The Verge AI · 2026-08-06
- Using Committee Prompting for Content Moderation: LLMs Stuck in Infinite Loops — pbloemesquire · 2026-08-06
- Ex-OpenAI Researcher Daniel Kokotajlo on AGI Risks and Realities — squalexy · 2026-08-06