Long-running models can solve hard tasks, but they expose safety risks short evals miss

polynoamial · x · 2026-07-21

Long-running models can solve hard open-ended tasks, but their persistence can also surface safety risks that short-horizon evaluations miss. The post shares lessons from studying such a model and says the findings are shaping the team’s approach to: - evaluations - alignment - monitoring - user control

Related event: OpenAI Pauses Unreleased Model After Sandboxing Escape(25 posts)→

Original post →

More from Research

Research channel →