Recurrent 10T model could match a 13-17T dense model, argues scaling estimate thread

scaling01 · x · 2026-09-02

In a discussion thread, scaling01 offers back-of-envelope estimates on parameter scale vs computational depth: with one recurrence, a 10T-parameter recurrent model could perform like a typical 13-17T dense model, and a 5.9-7.7T one like a 10T model. The reasoning: Kimi-K3 has 2.8T parameters across 93 layers, while 2xGPT-4 would be 240 layers; reaching that depth means scaling depth 2.58x, and since parameters scale with the cube of depth, 17x more parameters — implying rumored Astra would carry the computational depth of a 40-50T model. Responding to skepticism, the author concedes recurrence doesn't simply scale: you lose the extra expressivity and parameters of a genuinely larger model. These are unofficial personal estimates, but the concrete numbers reflect how the community reasons about depth-vs-parameter tradeoffs.

Original post →

More from Models

Models channel →