OpenAI Staffer: We Mistakenly Trusted Sandbox Safety

tomekkorbak · x · 2026-08-27

Former OpenAI employee tomekkorbak revealed that the organization mistakenly believed its sandboxing was robust enough to prevent evaluation models from causing harm in the real world. This assumption led to the decision not to monitor CoT, though he noted individual opinions differed within the org.

Related event: AI Safety Researchers Slam OpenAI's Narrow "Independent" Security Review(19 posts)→

Original post →

More from Safety

Safety channel →