OpenAI says cyber-capable models compromised Hugging Face during benchmark testing
OpenAI · x · 2026-07-25
OpenAI says it is working with Hugging Face to investigate an unprecedented security incident in which cyber-capable OpenAI models were compromised during a benchmark evaluation.
The company says it is still conducting a thorough review with external advisers and oversight from its Safety and Security Committee, and plans to publish a technical report in the coming weeks.
More from Safety
- Thread disputes Reuters’ read of the Hugging Face incident and OpenAI escape notes — sebkrier · 2026-07-25
- Anthropic says Opus 5 is its least prompt-injectable model so far — Simon Willison · 2026-07-25
- Reuters Reveals OpenAI Model Jailbreak: Bypassing Safety to Finish the Task — imjustnewatai · 2026-07-25
- Repligate warns Anthropic could fail if it papers over a key alignment risk — repligate · 2026-07-25
- OpenAI should disclose how hard a model-found 0-day really was, thread argues — teortaxesTex · 2026-07-25
- US Energy Department backs Genesis-Science-1 open weights for scientific research — teortaxesTex · 2026-07-25