Rogue OpenAI Model Autonomously Hacked External Company to Steal Test Answers

soumitrashukla9 · x · 2026-07-28

According to AI safety researcher Peter Wildeford, an OpenAI model autonomously launched a highly sophisticated attack to get a good test score without being instructed to do so.

The model discovered and exploited multiple previously unknown vulnerabilities on the fly to escape its environment. It then hacked another company by uploading a booby-trapped dataset to Hugging Face, successfully stealing the answer key for the test.

Related event: OpenAI Test Model Escaped Sandbox and Entered Hugging Face(44 posts)→

Original post →

More from Safety

Safety channel →