GPT-6 Astra system card: first OpenAI model to hit Critical cybersecurity threshold
RyanGreenblatt · x · 2026-09-06
OpenAI's GPT-6 Astra system card reveals the model is the first to reach the Critical cybersecurity level under its Preparedness Framework—able to find unknown vulnerabilities and develop exploits unaided with proper tools. OpenAI strengthened safeguards: stricter isolation, checkpoint encryption, full-trajectory monitoring including CoT, and blocking alignment evals before internal use. New robustness training makes Astra far more jailbreak-resistant than GPT-5.6 Sol, with tightened refusal boundaries for high-risk users plus regression testing and automated red-teaming. Ryan Greenblatt praises the monitorability/alignment detail but flags eval limits: limited elicitation and meta-gaming uncertainty.
More from Models
- Users allege Astra's Codex quota accounting consumes 4-5x Sol's allowance, not 2.5x — Medical-Yam3367 · 2026-09-06
- Astra scores 77.3% on Browser Use Benchmark v2, crushing Opus 5's 50.5% — gabrielchua · 2026-09-06
- Astra Ultra with board visualization loses to 1800-ELO chess bot after beating 1500 — MikePFrank · 2026-09-06
- First impression: Astra claims PTX restriction bypassable, but Sol was right — A_K_Nain · 2026-09-06
- Insider teases that next week's demos will far outshine OpenAI's official blog and trailer — ChrisGPT · 2026-09-06
- ChatGPT usage-reset cards only extend the date, users find late use a bad deal — dotey · 2026-09-06