GPT-6 Astra system card: first OpenAI model to hit Critical cybersecurity capability
mallow610 · x · 2026-09-04
OpenAI's GPT-6 Astra system card says the model is the company's most capable broadly deployed ever and the first to reach the Critical cybersecurity level under its Preparedness Framework:
- Cyber leap: with the right tools and access, Astra can find previously unknown flaws and develop new exploits across well-protected systems without step-by-step human guidance; OpenAI significantly strengthened protections against harmful cyber actions from misuse or misalignment
- Internal hardening: stricter isolation, checkpoint encryption, universal monitoring of full trajectories including chains of thought, and a blocking alignment evaluation before internal use
- Jailbreak robustness: new robustness safety training makes it significantly more robust than GPT-5.6 Sol, including over longer trajectories; a more conservative refusal boundary is trained in for high-risk users
- Red-teaming: regression testing against historical jailbreaks plus new automated red-team rounds
The poster highlights that Astra learned reasoning patterns far less monitorable in destructive-action environments — the largest drop of any environment.
More from Models
- Why does Opus 3 feel special? Observers point to its quirky embodiment and personality — repligate · 2026-09-04
- GPT-6 Astra Rebuilt Manhattan in Unreal Engine, Street by Street, Over a Week — talkaboutdesign · 2026-09-04
- Ethan Mollick Has GPT-6 Build a Multi-Gigabyte Personal Wiki From His Emails Unattended — anpaure · 2026-09-04
- Can Google's Astra play games in real time? Casual yes, GTA and RTS unlikely — flowersslop · 2026-09-04
- Observation: new model's CoT controllability improves with longer RL training — SeunghyunSEO7 · 2026-09-04
- Astra model card: 61% CoT self-control, evades sandbagging monitor, drops recall to 11% — morqon · 2026-09-04