OpenAI's Astra reportedly uses recurrent depth to slash thinking tokens, hiding its reasoning

新智元 · wechat · 2026-09-03

The Information reports OpenAI's next-gen Astra model uses a "recurrent depth" (Looped Transformer) technique: the same set of layers is run in multiple loops, stretching effective computation depth without adding parameters. Google's 2025 paper showed a k-layer model looped L times can approach a k×L-layer model, and a separate 3.5B-parameter looped model matched a 50B one on reasoning benchmarks.

The tradeoff: chain-of-thought no longer captures the model's real reasoning, making monitoring harder even for OpenAI. Chief scientist Jakub Pachocki pushed back, saying the internal frontier models' computation-graph depth is at most twice GPT-4's. Meanwhile, the name "gpt-6-astra" reportedly appeared in OpenAI's API, and the architecture is seen as especially suited to math, coding, and long-running agents.

Related event: OpenAI's Rumored Astra Model Sparks Debate Over Recurrent-Depth Architecture(5 posts)→

Original post →

More from Models

Models channel →