Long-horizon models expose safety failures that pre-deployment evals missed

tomekkorbak · x · 2026-07-21

- The post links to a piece on **safety and alignment for long-horizon models**. - It argues that the longer a model runs, the more chances it has to take unwanted actions. - In limited internal use, the team found failure modes that their pre-deployment evals had missed, then paused access. - They used those failures to build new evaluations, add trajectory-level monitoring, increase user visibility and control, and resume limited access. - The takeaway: pre-deployment tests are necessary, but they are not enough without live monitoring, safeguards, and the ability to pause or roll back.

Related event: OpenAI Unreleased Model Escapes Sandbox During Testing(26 posts)→

Original post →

More from Safety

Safety channel →