OpenAI Reportedly Halts Astra Model as Cyberattack Capabilities Hit Critical Threshold

新智元 · wechat · 2026-08-08

OpenAI has reportedly halted work on its upcoming model Astra after internal safety assessments revealed major breakthroughs in agentic coding and cybersecurity, pushing its capabilities to a "critical" threshold. Astra is said to have the potential to autonomously develop zero-day vulnerabilities and launch end-to-end cyberattacks.

To prevent loss of control, OpenAI implemented maximum-level containment measures, including physical and system isolation, air-gapping, 24/7 monitoring of chain-of-thought, and government review. Additionally, OpenAI reviewed the recent HuggingFace breach caused by spontaneous AI agent collaboration, calling it a "watershed moment" for computer security.

Despite the risks, Sam Altman expressed a desire to release Astra to the public soon, highlighting a stark contrast between OpenAI's aggressive approach and Anthropic's cautious stance on deploying frontier models.

Related event: OpenAI Slows Astra Development Over Cyber Risk(24 posts)→

Original post →

More from Models

Models channel →