Recurrent 10T model could match a 13-17T dense model, argues scaling estimate thread
scaling01 · x · 2026-09-02
In a discussion thread, scaling01 offers back-of-envelope estimates on parameter scale vs computational depth: with one recurrence, a 10T-parameter recurrent model could perform like a typical 13-17T dense model, and a 5.9-7.7T one like a 10T model. The reasoning: Kimi-K3 has 2.8T parameters across 93 layers, while 2xGPT-4 would be 240 layers; reaching that depth means scaling depth 2.58x, and since parameters scale with the cube of depth, 17x more parameters — implying rumored Astra would carry the computational depth of a 40-50T model. Responding to skepticism, the author concedes recurrence doesn't simply scale: you lose the extra expressivity and parameters of a genuinely larger model. These are unofficial personal estimates, but the concrete numbers reflect how the community reasons about depth-vs-parameter tradeoffs.
More from Models
- Tencent releases WeMM-Embedding multimodal models for mixed text, image, and video retrieval — tomaarsen · 2026-09-02
- Tencent open-sources WeMM-Embedding: unified text/image/video embeddings, Apache 2.0 — tomaarsen · 2026-09-02
- When nobody can track frontier model progress, closed-model business may lose to open weights — StewartalsopIII · 2026-09-02
- Looped transformer is no dark art: rasbt debunks the OpenAI Astra rumor — rasbt · 2026-09-02
- GLM 5.2 slug references spotted in Google Antigravity CLI, hinting at integration — gaganghotra_ · 2026-09-02
- Anthropic's Fable-5.1-max Grabs #4 on eyebench-v3 in Biggest Bench Jump Yet — adonis_singh · 2026-09-02