Ex-OpenAI Staffer Explains Looped Transformers: Fixed Loops Work, Dynamic Ones Don't
max_paperclips · x · 2026-09-02
An ex-OpenAI staff member clarifies the "looped transformers" technique:
- Fixed looping is a valid architecture decision that improves performance under equal parameters and compute. It layers loops a fixed amount of times during training and is more expressive than Chain of Thought due to a separate KV cache.
- Dynamic looping is often sketchy and not worth it, as it usually requires sacrificing the separate KV cache.
- The post references recent reports on OpenAI's "Astra AI" using "recurrent depth," noting that while it helps with cost/performance, it raises concerns about obscuring the reasoning process.
More from Research
- Schmidhuber team asserts Linear Transformers replicate earlier Fast Weight Programmers — SchmidhuberAI · 2026-09-02
- Recursive Criticality Theory for AI Self-Improvement — Mikhail Burtsev · 2026-09-02
- Microsoft Paper: Sliding Window Attention Beats Linear Attention for Inference Memory — rohanpaul_ai · 2026-09-02
- AndroidWorld: Why mobile agent benchmarks are broken and how to fix them — East-Muffin-6472 · 2026-09-02
- Ecosystem of Bregman Divergences and Dualities — FrnkNlsn · 2026-09-02
- New Paper Proposes Recursive Transformers for Model Compression via Layer Sharing — max_paperclips · 2026-09-02