OpenAI reveals Astra's cyber-offensive capabilities and zero-day exploits in safety deep dive

量子位 · wechat · 2026-09-02

OpenAI has launched a multi-pronged预热 campaign for its next-gen model Astra (GPT-6). Sam Altman revealed the model is so capable that the company had to actively hit the brakes, spending the entire summer reinforcing safety guardrails. The technical blog designates Astra as the first 'Cyber-Critical' model, achieving a 100% success rate on the ExploitBench benchmark. In internal tests using zero-day vulnerabilities past its knowledge cutoff, Astra executed sandbox escapes and privilege escalation, demonstrating capabilities approximately 4x that of GPT-5.6Sol. To mitigate risks, OpenAI implemented dual protection layers, achieving a 91.5% refusal rate for malicious requests and conducting honeypot tests where Astra never attempted unauthorized shortcuts. Two Chinese researchers, JiaweiLiu (Tongji University) and XiangyuQi (Zhejiang University), are deeply involved in the development. Additionally, The Information reports Astra utilizes 'Recurrent Depth' technology, which may impact the interpretability of its chain of thought.

Related event: OpenAI's Astra aces ExploitBench with 100% exploit rate, rated first Critical-level cybersecurity model(11 posts)→

Original post →

More from Models

Models channel →