Back-of-envelope: 40T-param sparse models are inevitable, DeepSeek on track
teortaxesTex · x · 2026-09-17
A back-of-envelope projection argues that 40T total parameters with 1T active per token should suffice, and DeepSeek's sparsity route (10T with Engram) is on track. Supernodes make serving 40T params easy, and synthetic data plus multimodal can cover 200T tokens — so giant sparse models are 'inevitable'.
More from Infra
- LinkedIn to present a PyTorch-native GPU retrieval engine powering feed and search — PyTorch · 2026-09-18
- vLLM boosts Kimi K3 serving throughput 2.2-2.8x with scheduler, KDA and MoE kernel optimizations — vllm_project · 2026-09-17
- AWS compares Bedrock RAG vector stores: OpenSearch vs pgvector vs S3 Vectors — AWS ML Blog · 2026-09-17
- Baseten makes pyannote diarization 9.6x faster with 3.2x higher throughput — baseten · 2026-09-17
- huggingface_hub v1.32.0 ships shared download cache, reusing 16.8GB of weights across repos — vanstriendaniel · 2026-09-17
- Fujitsu formally announces MONAKA, its 144-core ARM datacenter CPU — Fcking_Chuck · 2026-09-17