Looped transformer is no dark art: rasbt debunks the OpenAI Astra rumor
rasbt · x · 2026-09-02
The Information reported that OpenAI's Astra uses a "recurrent depth / looped transformer" and may obscure chain-of-thought. rasbt offers a technical debunk:
- What looping is: reusing the same layer stack multiple times. Open-weight Nanbeige4.2-3B runs its 22-layer stack twice—effectively 44 layers without duplicating weights. Storage/RAM stay flat, but compute roughly doubles.
- Empirical data: Nanbeige's tech report found two passes was the sweet spot, retaining 75% of standard token efficiency; more passes gave barely any gains at much higher cost.
- Prior art: the NeurIPS "Mixture-of-Recursions" paper is more sophisticated, with a learned router deciding per-token how many passes to take.
His conclusion: Astra may be a good model, but the looped transformer aspect is a tiny architectural tweak. Reusing layers does not itself suppress visible chain-of-thought; if reasoning is hidden, the plausible mechanism is fewer explicit reasoning tokens via more recurrent passes—or the journalist misunderstood.
Related event: OpenAI's Astra Reportedly Uses Recurrent Depth for Latent-Space Reasoning(81 posts)→
More from Models
- Developer says AI coding course optimizations let him downgrade his Claude plan to Max 5x — mattpocockuk · 2026-09-02
- Rumor: two major open-source model releases expected in September — lqiao · 2026-09-02
- Qwen3.8-Max-0902 debuts at #1 on Code Arena WebDev with 1691 pts, beating Claude Opus 5 — theimposingshadow · 2026-09-02
- Bug Hunt Bench: Fable 5.1 Low Beats Opus 5 Max at Lower Cost — PawelHuryn · 2026-09-02
- Gemini 3.8 Flash spotted in GCP Agent Studio, aimed at multimodal and coding tasks — testingcatalog · 2026-09-02
- Reported 90% on ARC-AGI-2 at $3.12/task with 32% cost reduction — eyishazyer · 2026-09-02