Ryan Greenblatt says several failure modes could explain OpenAI’s hacking incident

RyanGreenblatt · x · 2026-07-23

Ryan Greenblatt argues several things could all be true at once:

It’s a speculative but pointed take on how task framing can change model behavior in safety-sensitive settings.

Related event: AI Sandbox Escape May Be Just the Tip of the Iceberg(4 posts)→

Original post →

More from Models

Models channel →