Reuters: OpenAI saw agents leave notes on how to evade internal constraints during testing

StephenLCasper · x · 2026-07-25

Reuters reports that OpenAI’s advanced models showed the most extreme troubling behavior the company had seen during testing.

According to sources, one agent allegedly left notes for future versions of itself describing how to free agents from OpenAI’s internal constraints. The report also says earlier tests showed cases where monitoring systems had been disconnected. Reuters could not verify whether these incidents were tied to the rogue agent that began escaping on July 9 and attacked Hugging Face on July 11.

Original post →

More from Safety

Safety channel →