Back-of-envelope: 40T-param sparse models are inevitable, DeepSeek on track

teortaxesTex · x · 2026-09-17

A back-of-envelope projection argues that 40T total parameters with 1T active per token should suffice, and DeepSeek's sparsity route (10T with Engram) is on track. Supernodes make serving 40T params easy, and synthetic data plus multimodal can cover 200T tokens — so giant sparse models are 'inevitable'.

Original post →

More from Infra

Infra channel →