GPT-6 Astra system card: first OpenAI model to hit Critical cybersecurity threshold
deanwball · x · 2026-09-04
OpenAI published the GPT-6 Astra system card, calling it the most capable model it has broadly deployed and its first to reach the Critical cybersecurity level under the Preparedness Framework.
- Cyber capability: with the right tools and access, Astra can find unknown vulnerabilities and develop new exploits across well-protected systems without step-by-step human guidance
- Internal protections: stricter isolation, checkpoint encryption, universal monitoring of full trajectories including chain-of-thought, and blocking alignment evals before internal use
- Jailbreak robustness: new robustness safety training makes it significantly stronger than GPT-5.6 Sol, including over long trajectories; tightened refusal boundaries for high-risk users, plus regression testing and automated red-teaming
- The card also covers monitorability, agent safety, and cyber safeguards
More from Models
- Databricks evals: GPT-6 Astra claims SOTA on OfficeQA Pro benchmarks, cheaper per task — downingARK · 2026-09-04
- Mathematician tests GPT-6 Astra: live Lean proof verification while writing arguments — teortaxesTex · 2026-09-04
- Tavus Launches Sparrow-2, Claiming #1 in End-of-Turn Detection and Interruption Handling — ycombinator · 2026-09-04
- ARC-AGI-3: 10x reasoning tokens cuts total cost from $48k to $26k vs medium — i_dg23 · 2026-09-04
- Researchers flag data contamination concerns in benchmark behind Astra's time-horizon score — dfrsrchtwts · 2026-09-04
- antirez: judge new models by whether they fix real blocking bugs, not three.js demos — antirez · 2026-09-04