Report: OpenAI's Astra uses recurrent depth and hits critical cyber threshold

rohanpaul_ai · x · 2026-09-02

The Information reports Astra uses "recurrent depth" (a looped transformer): the same information runs through the same transformer layers multiple times per token, giving more compute per token without proportionally more parameters — a smaller model behaving like a larger one with less memory and bandwidth. The concern: more reasoning happens in internal numerical states rather than readable chain-of-thought, making human monitoring harder.

OpenAI previously said Astra is its first model to reach the Critical cybersecurity capability threshold under its Preparedness Framework: with the right tools and access, it can find unknown flaws and develop exploits across hardened systems without step-by-step human guidance. In internal evals Astra hit 39% exploit success at 75K output tokens, vs GPT-5.6 Sol at 1% there and 12% only near 140K tokens. Astra will launch with additional chain-of-thought monitoring, and classifiers can automatically stop potentially unauthorized actions.

Related event: OpenAI's Astra reportedly uses recurrent depth architecture, hiding reasoning and raising safety concerns(12 posts)→

Original post →

More from Models

Models channel →