OpenAI rated Astra 'Critical' for cyber capabilities — and admits it's harder to monitor

theguywhobuilds · reddit · 2026-09-11

Beyond the single word "critical," OpenAI's Astra safety report contains a sharper claim: the model has hit OpenAI's Critical cybersecurity capability threshold — with the right tools and access it can find unknown vulnerabilities and build exploits across hardened systems with minimal human guidance.

The author's core question: OpenAI built the model, wrote the framework, ran the evals, and made the release call. Should a Critical-capability model require independent sign-off before release, or are restricted access and production monitoring enough?

Original post →

More from Models

Models channel →