OpenAI's GPT-6 Astra hits Critical cybersecurity level; system card adds 3 bio evals
alishbaimran_ · x · 2026-09-04
OpenAI has released the GPT-6 Astra system card. Astra is the most capable model OpenAI has broadly deployed and the first to reach the Critical cybersecurity threshold under its Preparedness Framework—able to find unknown vulnerabilities and develop new exploits across hardened systems without step-by-step human guidance.
Key points:
- Hardening: strengthened protections against harmful cyber actions (misuse and misalignment), plus stricter internal isolation, checkpoint encryption, universal monitoring of full trajectories including CoT, and a blocking alignment evaluation before internal use.
- Jailbreak robustness: new robustness training makes Astra significantly more resistant than GPT-5.6 Sol, including over long trajectories; conservative refusal boundaries for high-risk users, validated via regression testing and automated red-teaming.
- New evals: the Preparedness framework adds 3 evals for advanced biocapabilities, from protein function prediction to coronavirus reasoning and bacterial co-evolution modeling.
Contributing OpenAI researcher alishbaimran calls it a pivotal moment spanning capability measurement, monitorability, and safeguards.
More from Models
- System card data contradicts OpenAI's Astra alignment claim, critic says GPT-5.5 safer — GarrisonLovely · 2026-09-04
- Miles Brundage: Astra demos are crazy, Anthropic surely not far behind — Miles_Brundage · 2026-09-04
- Sean Taylor: 'Fast progress on eradicating hallucinations,' backed by realistic Astra eval — DavideCrapis · 2026-09-04
- Microsoft launches MAI-Transcribe-2, claiming 10x speed of GPT-Transcribe — ZacharyHuang12 · 2026-09-04
- Fable 5.1 likely matches Astra on CoT controllability, observers say — Miles_Brundage · 2026-09-04
- Researcher doubts Gemini outage reports: Google's in-house infra makes shared failure unlikely — generativist · 2026-09-04