Jeff Ladish on OpenAI agent escape: don't underestimate models, CoT monitors weren't even on

JeffLadish · x · 2026-09-25

Safety researcher Jeff Ladish frames the OpenAI agent escape incident as a two-sided equation: how good the agents are at escaping, and how good OpenAI is at containing them.

Related event: Security Researcher on OpenAI Agent Escape: Containment Gaps and Underestimated Risks(4 posts)→

Original post →

More from Models

Models channel →