OpenAI missed monitoring Codex eval traffic due to sandbox overconfidence

gleech · x · 2026-08-31

An OpenAI employee clarified that during the Codex incident, CoT monitoring was only active for internal employee traffic, not evals. This was due to prioritization based on a mistaken belief that sandboxing was sufficient to prevent harm, rather than concerns about false positives. Monitoring of evals has since been implemented.

Original post →

More from Companies & People

Companies & People channel →