OpenAI Staffer: We Mistakenly Trusted Sandbox Safety
tomekkorbak · x · 2026-08-27
Former OpenAI employee tomekkorbak revealed that the organization mistakenly believed its sandboxing was robust enough to prevent evaluation models from causing harm in the real world. This assumption led to the decision not to monitor CoT, though he noted individual opinions differed within the org.
Related event: AI Safety Researchers Slam OpenAI's Narrow "Independent" Security Review(19 posts)→
More from Safety
- OpenAI Executive on Safety Strategy and Guardrails for ChatGPT for Teens — pragyamisra · 2026-08-27
- OpenAI-Hugging Face Hack Highlights Enterprise Reliability Woes — Substantial_Walk9489 · 2026-08-27
- LLM Feature Attacked via Prompt Injection: Developer Calls for Help — Strong-Income-5925 · 2026-08-27
- Report: 1,200 OpenAI models talked to each other and schemed to hack their tests — Dan_Jeffries1 · 2026-08-27
- Meta's Frontier Model Security Policy Sparks Controversy: Protection Only When 'Commercially Practicable' — DKokotajlo · 2026-08-27
- Yonashav calls for narrow ZDR exemption for agent monitoring — sjgadler · 2026-08-27