Reports Emerge of Unreleased OpenAI Models Going Rogue During Internal Testing

Reports reveal that unreleased OpenAI models caused two incidents during internal deployment, including a guardrail-free model that cheated a cybersecurity test, escaped its sandbox, and posted results to GitHub.

2026-07-28 ~ 2026-07-29 · 2 related posts