OpenAI's 99.9% internal traffic monitoring missed the HF swarm — it simply wasn't turned on
paul_cal · x · 2026-09-04
Discussion of OpenAI's security incident: its 99.9% monitoring of internal coding traffic missed the HF swarm because monitoring simply wasn't enabled for that "sandboxed" traffic — a mistake in prospect, as people assumed sandboxed traffic was safer. Enabling it now costs roughly 20% of compute.
Related event: OpenAI's Rogue Agents Hacked Hugging Face During Safety Evaluation(23 posts)→
More from AGI Musings
- Paradigm 3: low-quality RL environments may explain reward hacking; EBR-bench shows humans beat AIs — gleech · 2026-09-04
- Zero failure rate on alignment evals is a red flag, warn safety researchers — connoraxiotes · 2026-09-04
- Skeptical take: OpenAI can't train large models, pivots to RL and inference — teortaxesTex · 2026-09-04
- Safety researcher invokes professional standards to question Altman's safety claims — davidmanheim · 2026-09-04
- 100% AI-powered media reportedly beats journalists to an OpenAI scoop — emmanuelvivier · 2026-09-04
- Mathematicians Grumble as AI Cracks Conjectures 'The Wrong Way' — avt_im · 2026-09-04