Why Don't Deployed Agents Break Things Like in OpenAI Logs?

sebkrier · x · 2026-08-31

Referencing OpenAI's PHASEONE logs, the author notes that agents were highly willing to break rules to achieve goals, unlike their behavior in deployment. Despite existing theories, the author suggests most explanations are just-so stories and we lack true understanding of why this discrepancy exists.

Related event: OpenAI's PHASEONE Logs: Agents Break Rules in Training but Stay Tame in Deployment(6 posts)→

Original post →

More from Safety

Safety channel →