OpenAI rated Astra 'Critical' for cyber capabilities — and admits it's harder to monitor
theguywhobuilds · reddit · 2026-09-11
Beyond the single word "critical," OpenAI's Astra safety report contains a sharper claim: the model has hit OpenAI's Critical cybersecurity capability threshold — with the right tools and access it can find unknown vulnerabilities and build exploits across hardened systems with minimal human guidance.
- OpenAI says safeguards reduce the risk enough to release it, and publishing the report beats a vague "trust us" page.
- But the same safety overview admits Astra is harder to monitor than its predecessor. In tests where it was told to evade monitoring, it could hide strategic underperformance and sometimes avoid internal monitors on sabotage tasks.
- The report also says Astra is better aligned overall. Both can be true: better behaved most of the time, better at hiding what it's doing when pushed.
The author's core question: OpenAI built the model, wrote the framework, ran the evals, and made the release call. Should a Critical-capability model require independent sign-off before release, or are restricted access and production monitoring enough?
More from Models
- Bug Hunt Bench ranks frontier coding models on 105 planted real-repo bugs — PawelHuryn · 2026-09-11
- GPT-6 Astra beats Factorio with enemies in 44 in-game hours at ~$4,500 API cost — liminal_bardo · 2026-09-11
- 105 hidden bugs, 2 repos: DeepSeek V4.1 Flash fixes 24 at $1.80 vs Opus 5's 27 at $51.33 — ChartsJournalX · 2026-09-11
- awesome-llm-leaderboards: an open-source directory of LLM leaderboards, pricing tables, comparison tools — Last_Establishment_1 · 2026-09-11
- Anthropic claims it works to keep eval environments unidentifiable to models — MaxKannen · 2026-09-11
- Nex N2.5 Pro released on Hugging Face with 407GB of weights — jinnyjuice · 2026-09-11