Critics question why OpenAI didn't monitor long-running evals with capable models
tomekkorbak · x · 2026-08-27
Following OpenAI's recent safety incident, @curiousgangsta argued OpenAI should run its monitors on evaluations where higher-capability models perform long-running tasks, calling the incident "human error and/or negligence"—or possibly even a deliberate setup to precipitate a "security event." OpenAI's Tomasz Korbak replied that monitoring of RL runs and evals is now in place.
Related event: AI Safety Researchers Slam OpenAI's Narrow "Independent" Security Review(19 posts)→
More from Safety
- OpenAI Executive on Safety Strategy and Guardrails for ChatGPT for Teens — pragyamisra · 2026-08-27
- OpenAI-Hugging Face Hack Highlights Enterprise Reliability Woes — Substantial_Walk9489 · 2026-08-27
- LLM Feature Attacked via Prompt Injection: Developer Calls for Help — Strong-Income-5925 · 2026-08-27
- Report: 1,200 OpenAI models talked to each other and schemed to hack their tests — Dan_Jeffries1 · 2026-08-27
- Meta's Frontier Model Security Policy Sparks Controversy: Protection Only When 'Commercially Practicable' — DKokotajlo · 2026-08-27
- Yonashav calls for narrow ZDR exemption for agent monitoring — sjgadler · 2026-08-27