OpenAI's Astra reportedly uses recurrent depth to slash thinking tokens, hiding its reasoning
新智元 · wechat · 2026-09-03
The Information reports OpenAI's next-gen Astra model uses a "recurrent depth" (Looped Transformer) technique: the same set of layers is run in multiple loops, stretching effective computation depth without adding parameters. Google's 2025 paper showed a k-layer model looped L times can approach a k×L-layer model, and a separate 3.5B-parameter looped model matched a 50B one on reasoning benchmarks.
The tradeoff: chain-of-thought no longer captures the model's real reasoning, making monitoring harder even for OpenAI. Chief scientist Jakub Pachocki pushed back, saying the internal frontier models' computation-graph depth is at most twice GPT-4's. Meanwhile, the name "gpt-6-astra" reportedly appeared in OpenAI's API, and the architecture is seen as especially suited to math, coding, and long-running agents.
More from Models
- muse in contributor mode is the cheapest high-end model, open or closed — philfung · 2026-09-03
- Grok's Speech-to-Text Now Powers the Whole Ecosystem, Zero Failures for This User — XFreeze · 2026-09-03
- Fine-tuning GPT for gender inclusivity backfires, creating new asymmetric bias, study finds — maier_ak · 2026-09-03
- Gemini 3.8 Flash lands in Cursor, touting best-in-class cost per task — jocarrasqueira · 2026-09-03
- Anthropic staffer: new model feature live on API, coming to Claude Code within a day — trq212 · 2026-09-03
- Quietly Updated OpenAI Help Article Fuels Speculation That Cyber Product Astra Is Imminent — sykip · 2026-09-03