Astra rumored to use Recirculation for 1.38x effective parameter boost
scaling01 · x · 2026-09-02
Discussion suggests Astra might use "Recirculation" or recurrent depth techniques. While pre-training compute and decode latency remain unchanged, inference FLOPs and prefill latency increase. A related recurrent depth paper found that running the same recurrent block twice provides a 1.38× effective parameter multiplier. This implies OpenAI could train a 10T recurrent depth model that performs like a 13.8T standard model.
More from Models
- OpenAI previews Astra: a cybersecurity model scoring 100% on ExploitBench — LingmingZhang · 2026-09-02
- Users report Claude Code system prompt upgrade with toned-down personality — ivan_bezdomny · 2026-09-02
- Users report DeepSeek V4 Pro giving irrelevant answers — gefei55 · 2026-09-02
- Running 104GB Qwen3.8-Flash-Next on 48GB Mac at ~12 tok/s — yogthos · 2026-09-02
- GLM-5.3 Hits 310 tok/s, Coding Performance Competes with Opus — Yuchenj_UW · 2026-09-02
- User cancels Claude Max over confusing rate limits and new restrictions — robleclerc · 2026-09-02