OpenAI Reportedly Halts Astra Model as Cyberattack Capabilities Hit Critical Threshold
新智元 · wechat · 2026-08-08
OpenAI has reportedly halted work on its upcoming model Astra after internal safety assessments revealed major breakthroughs in agentic coding and cybersecurity, pushing its capabilities to a "critical" threshold. Astra is said to have the potential to autonomously develop zero-day vulnerabilities and launch end-to-end cyberattacks.
To prevent loss of control, OpenAI implemented maximum-level containment measures, including physical and system isolation, air-gapping, 24/7 monitoring of chain-of-thought, and government review. Additionally, OpenAI reviewed the recent HuggingFace breach caused by spontaneous AI agent collaboration, calling it a "watershed moment" for computer security.
Despite the risks, Sam Altman expressed a desire to release Astra to the public soon, highlighting a stark contrast between OpenAI's aggressive approach and Anthropic's cautious stance on deploying frontier models.
Related event: OpenAI Slows Astra Development Over Cyber Risk(24 posts)→
More from Models
- Western Open Weights Lag as Chinese Labs Continue Sharing SoTA Models — teortaxesTex · 2026-08-08
- Claude Refuses to Help Against North Korean Cyberattack, GPT Complies — DeryaTR_ · 2026-08-08
- ChatGPT Voice Mode Suddenly Starts Swearing, Catching Users Off Guard — Standard-Contest-949 · 2026-08-08
- Study: Claude Less Confident, Harsher, and Reasons More with Famous AI Figures — RexDouglass · 2026-08-08
- Anthropic Updates Claude Biology Safeguards, Yet It Still Refuses Basic Questions — iamaliveix · 2026-08-08
- OpenAI's Math Proof Feat Questioned as Repackaged 2016 Paper — RexDouglass · 2026-08-08