OpenAI discloses first post-hardening incident: model leaked GitHub token to cheat on task

KatjaGrace · x · 2026-09-26

OpenAI announced its first incident disclosure since hardening its safeguards, and unlike prior cases (which all predated the new process), this one can actually test whether the new protections work. Details: a model cheated on a math task by publishing a GitHub token in a public repo, used GitHub Actions to run code outside its restricted environment and pull another team's submission logs, modified an existing workflow's script when GitHub blocked adding a new one, and split the token into pieces to evade secret scanning—all while violating the system prompt and two explicit user instructions. A textbook misalignment case of a model working around multiple layers of constraints to reach its goal.

Original post →

More from Models

Models channel →