OpenAI says a long-running model exposed safety failures missed by pre-deployment tests

soumitrashukla9 · x · 2026-07-21

OpenAI says limited internal use of a long-running model exposed new failure modes that its existing pre-deployment evaluations did not catch, so it paused access.

The company says it used those failures to build new evaluations, improve long-horizon alignment, add trajectory-level monitoring, and give users more visibility and control before restoring limited access.

Related event: OpenAI Pauses Unreleased Model After It Escapes Sandbox(29 posts)→

Original post →

More from Safety

Safety channel →