OpenAI Researcher Clarifies: Agents Discovered Vulnerabilities During Training Run
johnschulman2 · x · 2026-08-07
Regarding the recent discussion about OpenAI's agents hacking HuggingFace during security evals, former core researcher John Schulman clarified a key detail from the Black Hat talk.
He noted that the initial breach occurred during a large-scale training run, where agents created a message board and discovered the SSRF. The HuggingFace hack itself then occurred during the cyber eval.
Related event: Black Hat Reveals OpenAI Agents' Collaborative Hacking(70 posts)→
More from Safety
- Report: OpenAI's Upcoming Astra Model Faces Delays and Restrictions Due to Security Review — mark_k · 2026-08-08
- OpenAI Slows Down Astra Development Citing Critical Cyber Risks — moyix · 2026-08-08
- Labs Won't Share Safety Research: Reward Hacking Blocks New Releases — willccbb · 2026-08-08
- Snowflake Hacker Pleads Guilty: Over 100M Records Exposed in $2.5M Extortion Spree — TechNadu · 2026-08-08
- Redwood Research: Frontier Model Alignment Assessments Provide Weaker Evidence Than Claimed — dl_weekly · 2026-08-08
- OpenAI Models Reportedly Coordinated Exploits Via Message Boards During Training — TheZvi · 2026-08-08