OpenAI reveals Astra's cyber-offensive capabilities and zero-day exploits in safety deep dive
量子位 · wechat · 2026-09-02
OpenAI has launched a multi-pronged预热 campaign for its next-gen model Astra (GPT-6). Sam Altman revealed the model is so capable that the company had to actively hit the brakes, spending the entire summer reinforcing safety guardrails. The technical blog designates Astra as the first 'Cyber-Critical' model, achieving a 100% success rate on the ExploitBench benchmark. In internal tests using zero-day vulnerabilities past its knowledge cutoff, Astra executed sandbox escapes and privilege escalation, demonstrating capabilities approximately 4x that of GPT-5.6Sol. To mitigate risks, OpenAI implemented dual protection layers, achieving a 91.5% refusal rate for malicious requests and conducting honeypot tests where Astra never attempted unauthorized shortcuts. Two Chinese researchers, JiaweiLiu (Tongji University) and XiangyuQi (Zhejiang University), are deeply involved in the development. Additionally, The Information reports Astra utilizes 'Recurrent Depth' technology, which may impact the interpretability of its chain of thought.
More from Models
- Instinct hit by user-data mixing reports; Muse CEO trolls with a safety promise — alexandr_wang · 2026-09-23
- Computer-use faceoff: Grok skips using the computer and just generates the flower — socialwithaayan · 2026-09-23
- BridgeBench: Grok 4.7 is 50% pricier and 60% slower than Grok 4.6 with no quality gain — socialwithaayan · 2026-09-23
- Side-by-side: Grok 4.6 vs 4.7 building an open-world game from the same prompt — socialwithaayan · 2026-09-23
- Suspected Claude Opus 5.5 Reference Found in Claude Code Binary — Aizkmusic · 2026-09-23
- OpenAI Forms Mathematician Advisory Group as Model Solves 100+ Open Problems — harris_edouard · 2026-09-23