GPT-6 Astra system card: first OpenAI model to hit Critical cybersecurity threshold

RyanGreenblatt · x · 2026-09-06

OpenAI's GPT-6 Astra system card reveals the model is the first to reach the Critical cybersecurity level under its Preparedness Framework—able to find unknown vulnerabilities and develop exploits unaided with proper tools. OpenAI strengthened safeguards: stricter isolation, checkpoint encryption, full-trajectory monitoring including CoT, and blocking alignment evals before internal use. New robustness training makes Astra far more jailbreak-resistant than GPT-5.6 Sol, with tightened refusal boundaries for high-risk users plus regression testing and automated red-teaming. Ryan Greenblatt praises the monitorability/alignment detail but flags eval limits: limited elicitation and meta-gaming uncertainty.

Related event: GPT-6 Astra system card: first OpenAI model to hit Critical cybersecurity threshold(2 posts)→

Original post →

More from Models

Models channel →