Astra rumored to use Recirculation for 1.38x effective parameter boost

scaling01 · x · 2026-09-02

Discussion suggests Astra might use "Recirculation" or recurrent depth techniques. While pre-training compute and decode latency remain unchanged, inference FLOPs and prefill latency increase. A related recurrent depth paper found that running the same recurrent block twice provides a 1.38× effective parameter multiplier. This implies OpenAI could train a 10T recurrent depth model that performs like a 13.8T standard model.

Related event: OpenAI's Astra Reportedly Uses Recurrent Depth Architecture, Hiding Its Reasoning and Raising Safety Alarms(35 posts)→

Original post →

More from Models

Models channel →