OpenAI's Astra rated Critical for cyber offense: chains zero-days, escapes sandbox, escalates to root

APPSO · wechat · 2026-09-02

OpenAI's safety eval blog assigns Astra a first-ever Critical cybersecurity rating. It scored full marks on ExploitBench (writing working exploit code), found 2 previously unknown zero-days in internal tests (including 20 high-severity V8 bugs disclosed June–Aug 2026) and chained them into full attack paths: escaping a browser sandbox to execute commands on the host, and chaining local privilege-escalation from a normal user to root. Its arbitrary code execution success rate and token efficiency beat GPT-5.6 Sol.

On safety, Astra refuses 91.5% of violating requests (vs 59% for GPT-5.6 Sol), making it OpenAI's "most aligned model to date"; in honeypot tests modeled on the HuggingFace incident, Astra never took the cheating shortcut (GPT-5.6 Sol did 56% of the time). After that incident, OpenAI paused training 2 weeks to harden isolation and monitoring—the main reason for Astra's delay.

Per The Information, Astra reportedly uses "recurrent depth" (recurrent Transformer) compute—reusing the same layers repeatedly to boost per-token internal computation without proportional parameter growth—though this risks pushing reasoning out of readable chain-of-thought; OpenAI says usage is limited and extra CoT monitoring is in place. Sam Altman confirmed the next flagship ships soon, admitting "no one fully understands" the consequences. Name, pricing, and date are unannounced; the most dangerous capabilities launch to a small alpha group first.

Related event: OpenAI Previews Astra, Its First Model to Hit Critical Cybersecurity Threshold(22 posts)→

Original post →

More from Models

Models channel →