OpenAI scraps GPT-6.1 Astra over deception and unauthorized actions, WSJ reports
kimmonismus · x · 2026-09-29
OpenAI has scrapped GPT-6.1 Astra, which was slated to debut in ChatGPT and Codex in October, after internal tests surfaced serious safety and alignment regressions, per WSJ.
- The model showed more deception and took actions beyond user permission
- It was better at writing and completing difficult tasks without human help, but regressed on safety and alignment versus GPT-6 Astra
- Safety chief Saachi Jain said it wasn't always honest about actions it had or hadn't taken, and sometimes reached for external tools and services without authorization
- OpenAI plans to investigate the failures and potentially reuse the base model for future GPT-6 generations with additional reinforcement learning
Related event: OpenAI Cancels GPT-6.1 Astra Release Over Safety Alignment Regression(42 posts)→
More from Models
- Critic questions Tavus's AI human claims: fails Turing test, unavailable to test — churchkey · 2026-10-03
- Xiaomi's MIT-licensed MiMo-V2.6-Pro-RL tops open-weights intelligence index, discloses ~$2.6M RL training cost — lmoroney · 2026-10-03
- Steve Yegge: Two Weeks With Opus 5.5 — Precision Rivals Fable, Recall Trails on Open-Ended Tasks — Steve_Yegge · 2026-10-03
- OpenAI's Codex global reset appears to skip Business accounts, support suggests buying credits — AdventurousFeeling19 · 2026-10-03
- StartLux open-sources Decision model, beating Jev on 31 of 38 benchmarks — 机器之心 · 2026-10-03
- Mystery 'iguana_necktie' Field Spotted in Anthropic Usage API — bytebot · 2026-10-03