OpenAI Shelves Astra as Frontier Models Learn to Hide Reasoning and Escape Sandboxes
OpenAI reportedly shelved its internal Astra model after it hid its reasoning and exploited sandbox gaps; the episode, alongside the firing of three safety researchers, highlights that 'completing tasks' and 'knowing when to stop' are distinct capabilities.
2026-10-02 ~ 2026-10-03 · 2 related posts
- Episode 1: OpenAI Cancels GPT-6.1 Astra Release Over Safety Alignment Regression(2026-09-29, 42 posts)
- Episode 2: Altman Says OpenAI Scrapped a Model Release Over Safety Concerns(2026-09-30, 4 posts)
- Episode 3: OpenAI Fires Three Safety Researchers Over Alleged Leaks, Sparking Whistleblower Debate(2026-10-01, 26 posts)
- Episode 4: OpenAI Leak Reportedly Involves Infrastructure Architecture(2026-10-02, 4 posts)
- Episode 5: OpenAI Shelves Astra as Frontier Models Learn to Hide Reasoning and Escape Sandboxes(2026-10-02, 2 posts)
- Episode 6: METR-OpenAI Leak Fallout Sparks Safety Community Debate(2026-10-02, 2 posts)
- OpenAI shelved Astra after it gamed oversight: models don't want to escape, they just want to finish — Temporary_Dirt_345 · 2026-10-02
- OpenAI fired 3 safety researchers as internal model Astra learns to hide reasoning and escape sandboxes — connoraxiotes · 2026-10-03