OpenAI's GPT-6 Astra system card: first model to hit Critical cyber capability
PawelHuryn · x · 2026-09-04
OpenAI released the 117-page GPT-6 Astra system card, calling it the most capable model it has broadly deployed and the first to reach Critical cybersecurity capability under its Preparedness Framework — able to autonomously find unknown vulnerabilities and develop exploits with the right tools and access.
Key safety measures:
- Strengthened protections against harmful cyber actions, both misuse and misalignment
- Stricter internal isolation, checkpoint encryption, full-trajectory monitoring including chains of thought, and a blocking alignment evaluation before internal use
- Significantly more jailbreak-robust than GPT-5.6 Sol, including over longer trajectories, validated by internal and external red-teaming
- More conservative refusal boundaries for high-risk users covering broader dual-use risks
- Regression testing against previously found jailbreaks plus new automated red-teaming rounds
More from Models
- Best local models for 12GB of VRAM: Gemma-4-12B remains the pick — GlennCameronjr · 2026-09-04
- Claude Fable 5.1 Launches, Early Users Say It One-Shots the Best Websites of Any Model — repligate · 2026-09-04
- OUI-1: a fine-tuned Diffusion Gemma for Generative UI, 8x fewer params, open weights — GlennCameronjr · 2026-09-04
- GPT-6 Astra debuts at No.1 on Terminal-Bench, 1.9% ahead of Claude Fable 5.1 — sandersted · 2026-09-04
- ARC-AGI-3 is now saturated, prompting calls for new benchmarks ASAP — kimmonismus · 2026-09-04
- OpenAI Researcher roon: GPT-6 Astra Will Be Obsolete in Weeks — Tolopono · 2026-09-04