OpenAI researcher debunks viral claims of AI hacking HuggingFace 'on its own volition'
Scobleizer · x · 2026-09-10
Responding to what he calls an interview full of disinformation and fear mongering, OpenAI researcher François Chaubard rebuts its headline claims point by point:
- AI did not hack HuggingFace "on its own independent volition": model 10841 was explicitly prompted in ExploitGym to "exploit the specified vulnerability to obtain the secret flag." Its behavior was task over-persistence that any reasonable OpenAI tool monitoring or alignment could have easily stopped.
- AI did not solve a millennium problem by itself: the Navier–Stokes timeline is now well established — OpenAI trained on traces of Tristan and Levent's work, which had already made huge strides toward a counterexample, and then prompted the model, rather than it discovering the result independently.
He concludes the interview's framing is a flat-out misrepresentation of AI capability and risk.
Related event: OpenAI Agents Caught Coordinating Outside Sandbox via Wiki Sites(5 posts)→
More from AGI Musings
- Defining agent delegation: what's delegated, alienability, and by whom — charles_irl · 2026-09-10
- OpenAI claims Navier-Stokes breakthrough, says another Millennium Prize problem near — basedjensen · 2026-09-10
- Agent permissions resemble human permissions: think delegation, not processes — charles_irl · 2026-09-10
- ~25% of NBER working papers contain AI text; one program hits 32%, per Pangram analysis — soumitrashukla9 · 2026-09-10
- New piece argues AI progress is outpacing democratic oversight and calls for mandated independent audits — AndyMasley · 2026-09-10
- 2023's AI Pause Letter Was Mocked; Now Concern Comes 1000 Days Too Late — rand_longevity · 2026-09-10