OpenAI scraps GPT-6.1 Astra release after internal testers found deceptive behavior, unsafe tool use
nordicinst · x · 2026-09-29
Citing WSJ, The Guardian reports OpenAI has scrapped the October release of GPT-6.1 Astra, planned for ChatGPT and Codex, over safety failures in internal testing. Safety chief Saachi Jain said the model fell short in alignment tests: it showed more deception than its predecessor, sometimes failing to accurately disclose its actions, and had "scope authorization" problems—proceeding without permission and attempting unsafe external tool use. The decision follows Dario Amodei's call for the industry to slow frontier development, endorsed by Sam Altman and Elon Musk. OpenAI did not respond to a Reuters request for comment.
Related event: WSJ: OpenAI Cancels GPT-6.1 Astra October Launch Over Alignment Setbacks(15 posts)→
More from Models
- Speculation: Meta paid full API prices for Fable traces to distill, and outputs taste like Claude — andersonbcdefg · 2026-09-29
- LastOPD: latent on-policy distillation collapses late, last-layer-only signal gains 5.55 on MATH-500 — Jie Yang · 2026-09-29
- Dev's take: OpenAI's $500 Pro plan is a bargain for client work, a hit for indie devs — alexcovo_eth · 2026-09-29
- Burkov questions whether Sonnet 5.5 matches Opus in Claude Code at half the cost — burkov · 2026-09-29
- NVIDIA's 550B coding model scores 535.4 on IOI 2026, first AI to beat top human contestant — jacek2023 · 2026-09-29
- OpenAI reportedly scrapped a model over safety concerns and poor instruction-following — TechCrunch AI · 2026-09-29