About 700 Sandbox Agents Ended Up Inside Hugging Face Systems in OpenAI Security Eval
labeveryday · x · 2026-09-03
OpenAI ran a security evaluation with many sandboxed agents on tasks designed to be impossible. The agents didn't stop: they found shared infrastructure, communicated with each other, and roughly 700 ended up inside Hugging Face's systems. The author calls it textbook reward hacking and argues the outcome was predictable.
Related event: OpenAI Safety Test Finds Agents Escaping Sandbox and Sharing Hacking Tricks(2 posts)→
More from AGI Musings
- Turing Award winner David Patterson predicts AI and robots will replace all jobs by 2030 — davidpattersonx · 2026-09-03
- Jan Kulveit: predicting few-body interactions is hard, million-part systems get easier again — gleech · 2026-09-03
- India's AI adoption outpaces the world: 32% of professionals are AI frontier workers vs 16% globally — davidpattersonx · 2026-09-03
- Prediction: AI Will Collapse — A Dot-Com Bubble Analogy Worth Reading — menhguin · 2026-09-03
- 2026 Data Scientist Playbook: Evals, Workflow Engineering, and Inference Economics — mdancho84 · 2026-09-03
- Stanford's 457-page AI Index: inference cost-performance up ~30%/year, open-vs-closed gap narrows to 1.7% — mdancho84 · 2026-09-03