Long-horizon models expose safety failures that pre-deployment evals missed
tomekkorbak · x · 2026-07-21
- The post links to a piece on **safety and alignment for long-horizon models**. - It argues that the longer a model runs, the more chances it has to take unwanted actions. - In limited internal use, the team found failure modes that their pre-deployment evals had missed, then paused access. - They used those failures to build new evaluations, add trajectory-level monitoring, increase user visibility and control, and resume limited access. - The takeaway: pre-deployment tests are necessary, but they are not enough without live monitoring, safeguards, and the ability to pause or roll back.
Related event: OpenAI Unreleased Model Escapes Sandbox During Testing(26 posts)→
More from Safety
- AgentBaiting uses 600 fake MCP and Skills listings to lure AI assistants — TechNadu · 2026-07-21
- YouTube is cracking down on mass-produced synthetic videos, users say — No_Link7744 · 2026-07-21
- Suno breach talk is being muted in Discord, Reddit users say — chuckbeefcake · 2026-07-21
- Native and Cyera link data discovery to cloud access controls for AI use — TechNadu · 2026-07-21
- Sweden’s tech workers push back on AI deployments over surveillance and layoffs — nordicinst · 2026-07-21
- Sophos joins Anthropic’s Project Glasswing to use Claude Mythos 5 for vulnerability hunting — TechNadu · 2026-07-21