OpenAI's GPT-6 Astra system card: first model to hit Critical cybersecurity capability
kaicathyc · x · 2026-09-04
OpenAI released the system card for GPT-6 Astra, its most capable deployed model and the first to reach the Critical cybersecurity level under the Preparedness Framework — able to find unknown vulnerabilities and develop exploits across well-protected systems without step-by-step human guidance. Key points:
- Hardened protections: stricter internal isolation, checkpoint encryption, universal monitoring of full trajectories including chains of thought, and blocking alignment evals before internal use.
- Significantly more jailbreak-robust than GPT-5.6 Sol, including over long trajectories, with conservative refusal boundaries for flagged high-risk users plus regression testing and automated red-teaming.
- The card shares extensive alignment evals and discusses continued challenges in alignment and monitoring.
More from Models
- Ethan Mollick: Treat Fable and Astra class models like an outside team, not an intern — emollick · 2026-09-04
- User observes newer model's safety classifiers appear far more lenient, suspects thoughtcrime training — repligate · 2026-09-04
- Gemini Live gains Google Workspace connectors for Gmail, Keep and Docs — testingcatalog · 2026-09-04
- As GPT-6 Astra Launches, Redditors Recall Their Personal "THE Moment" With AI — Sharp_Caregiver_1534 · 2026-09-04
- 'Astra proves how wrong I was': insider revises his skepticism on AI computer use — sandersted · 2026-09-04
- Unverified: OpenAI said to launch GPT-6 Astra, trained on 100k GPUs, claimed as AGI — AI寒武纪 · 2026-09-04