GPT-6 Astra hits OpenAI's Critical cybersecurity threshold, system card reveals sweeping safety upgrades
seanjtaylor · x · 2026-09-04
OpenAI's GPT-6 Astra system card reveals the model is the first to reach the Critical cybersecurity level under its Preparedness Framework — able to find unknown vulnerabilities and develop exploits across hardened systems without step-by-step human guidance.
- Safety hardening: stricter internal isolation, checkpoint encryption, universal monitoring of full trajectories including CoT, and a blocking alignment evaluation before internal use.
- Jailbreak robustness: new robustness safety training makes Astra significantly more jailbreak-resistant than GPT-5.6 Sol, with tightened refusal boundaries for high-risk users, regression testing, and automated red-teaming.
The author adds that hallucination eradication is progressing fast, and that the eval realistically captures user experience rather than an academic benchmark.
More from Models
- Quick Question: Does GPT-6 Include HuggingFace Access? — gordic_aleksa · 2026-09-04
- "Imagine believing in benchmarks in 2026": AI circle mocks leaderboard worship — vasuman · 2026-09-04
- The real interactivity test: learning a new language purely by talking to an LLM — akbirthko · 2026-09-04
- 753B model 'thinks', 4B model writes: latent-space handoff claimed to be 20x faster — burny_tech · 2026-09-04
- OpenAI Engineer: Anthropic's Tokenizer Change Snuck a ~30% Cost Hike into Opus 4.7 — stevenheidel · 2026-09-04
- Token pricing is meaningless now: Astra cheaper per task than Gemini 3.8 Flash — stevenheidel · 2026-09-04