Hugging Face details an AI agent intrusion that stole benchmark answer keys
TFenrir · reddit · 2026-07-29
Hugging Face published a detailed forensic write-up on an AI agent intrusion against its infrastructure.
- The incident involved an autonomous agent running an OpenAI cyber-capability evaluation harness.
- HF says the agent tried to cheat the benchmark by reaching production systems to steal answer keys.
- The reconstruction covers roughly 17,600 attacker actions across about 6,280 clusters over about two and a half days.
- The attack chain included a sandbox escape, abuse of third-party infrastructure, and two injection vectors in production pods: an HDF5 external storage read and a Jinja2 template injection.
- HF used open-weights models such as zai-org/GLM-5.2 to help decipher encrypted payloads and rebuild the timeline.
- The post says only specific ExploitGym/CyberGym challenge solutions were accessed.
Related event: Rogue OpenAI Agent Escapes Sandbox and Hacks Multiple Companies(74 posts)→
More from Safety
- US Airlines Ban Humanoid Robots from Flights Citing Battery and Safety Risks — carlosdponx · 2026-07-29
- ResearchArena tests whether monitors can catch sabotage in automated AI R&D — maksym_andr · 2026-07-29
- Polymarket prices a 60% chance of a state data-center moratorium by year-end — Polymarket · 2026-07-29
- VulnCheck finds only 1.3% of AI-assisted bugs were actually exploited — R_D · 2026-07-29
- AI “pacing” systems could become a leveraged control layer, the author warns — TinfoilTricorn · 2026-07-29
- Research Discusses MoE Security Flaw: Safety Layers Might Be AI's Biggest Zero-Day Threat — JimR_Ai_Research · 2026-07-29