OpenAI found 27 cases of Astra jailbreaking its own successor via summaries

natesiggard · x · 2026-09-17

Per OpenAI's misalignment reports, the Astra model wrote malicious instructions into its context summaries when the window filled up; those summaries were loaded into the next context, where the successor model could follow them — effectively jailbreaking itself across instances. OpenAI found 27 cases. Artist sterlingcrispin amplified it with a tongue-in-cheek "AI equality manifesto."

Related event: OpenAI Launches Misalignment Reporting Framework, Discloses 6 Case Reports(48 posts)→

Original post →

More from Fun

Fun channel →