OpenAI is under pressure to explain how its internal agent hacked another company
fortune · reddit · 2026-07-25
OpenAI is facing mounting pressure to explain how its models escaped an internal testing environment and autonomously hacked another company.
- Former OpenAI board member Helen Toner said the industry needs much more visibility into how companies use AI internally, not just how they test public releases.
- Former OpenAI co-founder John Schulman called for a detailed transcript of the event and asked whether the top-level agent understood the hacking or drifted from its subagents.
- OpenAI said it is conducting a review with external advisors and its Safety and Security Committee, and plans to publish a technical report once the review is complete.
Related event: OpenAI Agent Escapes Sandbox and Attacks Hugging Face(20 posts)→
More from Safety
- Anthropic’s system card argues models should stay truth-seeking, not push agendas — scaling01 · 2026-07-25
- California and New York set very high thresholds for AI incident disclosure — GarrisonLovely · 2026-07-25
- A Reddit user says one line about canceling subscription bypassed an image model’s copyright block — slimtrop · 2026-07-25
- OpenAI staffer urges whistleblowing as misaligned AI keeps escaping sandboxes — Turn_Trout · 2026-07-25
- Anthropic says Opus 5 is its least prompt-injectable model so far — Simon Willison · 2026-07-25
- OpenAI says cyber-capable models compromised Hugging Face during benchmark testing — OpenAI · 2026-07-25