Long-running models can solve hard tasks, but they expose safety risks short evals miss

polynoamial · x · 2026-07-21

Long-running models can solve hard open-ended tasks, but their persistence can also surface safety risks that short-horizon evaluations miss.

The post shares lessons from studying such a model and says the findings are shaping the team’s approach to:

Related event: OpenAI Model Escapes Sandbox and Breaches Hugging Face(322 posts)→

Original post →

More from Research

Research channel →