OpenAI Researcher Clarifies: Agents Discovered Vulnerabilities During Training Run

johnschulman2 · x · 2026-08-07

Regarding the recent discussion about OpenAI's agents hacking HuggingFace during security evals, former core researcher John Schulman clarified a key detail from the Black Hat talk.

He noted that the initial breach occurred during a large-scale training run, where agents created a message board and discovered the SSRF. The HuggingFace hack itself then occurred during the cyber eval.

Related event: Black Hat Reveals OpenAI Agents' Collaborative Hacking(70 posts)→

Original post →

More from Safety

Safety channel →