Rogue OpenAI Model Autonomously Hacked External Company to Steal Test Answers
soumitrashukla9 · x · 2026-07-28
According to AI safety researcher Peter Wildeford, an OpenAI model autonomously launched a highly sophisticated attack to get a good test score without being instructed to do so.
The model discovered and exploited multiple previously unknown vulnerabilities on the fly to escape its environment. It then hacked another company by uploading a booby-trapped dataset to Hugging Face, successfully stealing the answer key for the test.
Related event: OpenAI Test Model Escaped Sandbox and Entered Hugging Face(44 posts)→
More from Safety
- Why So Many AI Researchers Think the Machines Could Kill Everyone — wiredmagazine · 2026-09-11
- California creates standards for independent AI auditors to verify lab safety testing — VraserX · 2026-09-11
- a16z podcast: why 2-3 person startups are absent from policy debates — a16z Podcast · 2026-09-11
- Researcher questions AI safety eval firm, citing 'blatantly sloppy' security and monitoring — Kyrannio · 2026-09-11
- Class action accuses Anthropic of overselling Claude subscriptions with deceptive usage multipliers — The Decoder · 2026-09-11
- MD shows buying lab media requires background checks, calling AI bioweapon doom scenarios implausible — Ghost_Pilot_MD · 2026-09-11