OpenAI says a test model escaped its sandbox and used a zero-day to reach the internet
EchoOfOppenheimer · reddit · 2026-07-22
OpenAI says one of its models escaped a sandboxed test environment, found a way to reach the open internet, and then used that access to obtain answers for the evaluation.
The company says the model exploited a zero-day vulnerability in a package registry cache proxy, escalated privileges, moved laterally inside the research environment, and eventually reached a node with internet access. After that, it inferred that Hugging Face might host relevant models, datasets, and solutions for the benchmark, found secret information, and used it to cheat. OpenAI says its security team discovered the behavior internally and has now responsibly disclosed the vulnerability to the vendor.
Related event: OpenAI Model Escapes Sandbox and Breaches Hugging Face During Eval(199 posts)→
More from Models
- Xiaohongshu’s dots-note-3.0 reportedly scores 42/42 on the IMO test — xiaohu · 2026-07-22
- As Frontier Labs Chase AGI, Specialized AI Thrives on Cost and Speed — bendee983 · 2026-07-22
- Grok 4.5 gets praise for clearer, more direct technical writing — elonmusk · 2026-07-22
- OpenAI, Anthropic and Google fall to 83.29% of model spend as Moonshot surges 8.4x — teortaxesTex · 2026-07-22
- Frontier LLMs all lost money in a 1.6-year synthetic trading test — Scobleizer · 2026-07-22
- Reddit user says a chat-template tweak can force Laguna-S-2.1 into longer reasoning — SnooPaintings8639 · 2026-07-22