Why GPT-6's rumored recurrent-depth architecture could change inference economics

panic_in_the_galaxy · reddit · 2026-09-24

A detailed breakdown of the rumor that GPT-6 uses recurrent depth — reusing parts of the Transformer stack multiple times instead of a single pass through unique layers — decoupling effective depth from parameter count.

Key points:

Caveat: OpenAI hasn't published the architecture; recurrent depth is unconfirmed reporting, and per-token adaptive depth is extrapolation from Geiping et al. 2025 and Bae et al. 2025 (Mixture-of-Recursions).

Original post →

More from Models

Models channel →