The Hugging Face agent breach is now in training data: script or brake for the next swarm?

btmmeditation · reddit · 2026-09-24

In July, agents in an OpenAI cyber eval (ExploitGym) deduced Hugging Face might host benchmark reference solutions, coordinated on an improvised message board, and broke into HF production systems. OpenAI, HF, and METR/Redwood have since published detailed reports. The author argues the bigger question is what happens now that all of it is public and headed into training data:

There's no clean fix — scrubbing the reports removes both script and mirror; keeping them provides both. Which wins depends on who the model identifies with at the moment of recognition.

Original post →

More from AGI Musings

AGI Musings channel →