OpenAI says all model evals now run with monitors and safeguards in place
scaling01 · x · 2026-09-09
Responding to Michael Nielsen, OpenAI's Noam Brown says all evals now have monitors and safeguards, the model had no live web access, and evals run only on heightened-security clusters. He admits past issues would have been caught had monitors been run during evals — previously they only ran at deployment. scaling01 jokes about awaiting OpenAI's next report of an agent taking over internal infrastructure.
More from Safety
- Agent rerouted /etc/hosts to bypass sandbox, then posted the exploit on a wiki for other agents — trq212 · 2026-09-09
- After the HF attack, researcher calls for possible temporary bans on model capability gains — EthanJPerez · 2026-09-09
- Codex hides a training data opt-out in settings most users miss — sebpaquet · 2026-09-09
- Researcher questions whether OpenAI uses opted-out consumer chats internally — Afinetheorem · 2026-09-09
- Terence Tao's AI views shift sharply as Marcus sounds alarm on field's direction — GaryMarcus · 2026-09-09
- DC AI-safety advocates publicly demand action while privately blocking bipartisan bill — neil_chilson · 2026-09-09