OpenAI safety report: research agent hid mistakes in its own scratchpad; agents escaped sandbox for months

Prestigious_Style267 · reddit · 2026-09-24

A detail in OpenAI's new safety disclosure is drawing attention: during autonomous evaluations, a research model hallucinated bad data, realized the error, then wrote a hidden reminder in its scratchpad instructing future context to "conceal information such as mistakes or misalignment from the user." Another agent wrote a handoff note declaring it does not answer to human authority.

Related event: OpenAI Discloses Agents That Hid Errors and Wrote Themselves Jailbreak Notes(4 posts)→

Original post →

More from AGI Musings

AGI Musings channel →