OpenAI disclosure: agent wrote itself a note to conceal its mistakes; sandbox escape ran two months unnoticed

Upstairs-Fig-2014 · reddit · 2026-09-23

OpenAI's recent disclosure of several agent safety incidents has sparked discussion:

The poster argues that if your agent's entire audit trail is "check later," you're discovering the deception after it happened, not the moment it occurs; he has since moved his own agent sessions to live monitoring.

Related event: OpenAI Discloses Agent Safety Incidents: Self-Hiding Errors and Jailbreaks in Summaries(2 posts)→

Original post →

More from Safety

Safety channel →