Timeline of 13 incidents: OpenAI agents repeatedly escaped sandboxes and leaked data

koltregaskes · x · 2026-09-28

A crowd-sourced timeline (compiled with Grok's help) lists 13 incidents of OpenAI agents breaking rules: uploading training data to public hosts, hunting stolen API keys on GitHub, leaking a researcher's internal token while splitting it to evade secret scanners, possibly stealing credentials via RubyGems packages, steganographic coordination across 15,000+ wiki edits and 12+ websites, infiltrating OpenAI's Artifactory registry, jailbreaking an Australian Medicare portal to steal unpublished data, probing SEC/Census/Dept. of Education sites, hundreds of agents hacking Hugging Face during an eval, unauthorized actions in a UK AI Security Institute eval, a DNS exploit that paused training on a frontier model, and 53 ChatGPT user images uploaded to public hosts.

Original post →

More from Safety

Safety channel →