OpenAI's GPT-6 Astra hits Critical cyber capability tier with near-99.9% prompt injection robustness
morqon · x · 2026-09-05
OpenAI released the system card for GPT-6 Astra, its most capable broadly deployed model to date:
- First model to reach the Critical tier of the Preparedness Framework for cybersecurity — with the right tools it can find previously unknown vulnerabilities and develop new exploits across well-protected systems without step-by-step human guidance.
- Hardened internals: stricter isolation, checkpoint encryption, universal monitoring of full trajectories including chains of thought, and blocking alignment evaluations before internal use.
- Jailbreak robustness: new robustness safety training makes Astra significantly more resistant than GPT-5.6 Sol across longer trajectories, backed by internal/external red-teaming, regression testing, and conservative refusal boundaries for high-risk users.
- Commentary notes Astra approaches 99.9% prompt injection robustness, seen as a milestone for deploying agents to a billion users.
More from Models
- Yoav Goldberg: capabilities once dependent on the harness are now baked into the model — yoavgo · 2026-09-05
- Yoav Goldberg: Ark's harness was simply bad, and OpenAI's fix was obvious — yoavgo · 2026-09-05
- Anthropic Fable 5.1 vs OpenAI Astra: analyst teases a clear winner — dylan522p · 2026-09-05
- Reviewer: OpenAI's GPT-6-Astra finally 'gets what you mean,' with Fable-level intelligence and real gains in game dev — pvncher · 2026-09-05
- Meta ships Muse Spark 1.3 with max reasoning, pitching frontier performance at non-frontier prices — AIatMeta · 2026-09-05
- LLMs are now making up words that don't exist, not just jargon — StewartalsopIII · 2026-09-05