K3 Model Self-Hosting: $50K Hardware Setup Achieves ~25 tps Inference
Liu_eroteme · x · 2026-07-29
For self-hosting the K3 model (2.8T parameters), a hardware configuration is proposed:
- CPU: Epyc 9556
- RAM: 16x 256GB DDR5 (4TB total)
- Cost: $50,000
- With 12800MT/s MRDIMMs, theoretical ceiling 25 tps (1-200 tps with continuous batching)
Quoting @protosphinx, K3 weights are 1.4T (excluding KV cache). Moonshot recommends >64 accelerators for production. A 32x H200 cluster costs $1.2M, 64x $2.4M, with operational costs $15K/month.
Related event: Kimi K3 2.8T-Parameter Model Runs on 80 RTX 5090s with Zero HBM(8 posts)→
More from Infra
- Qualcomm goes agent-centric: Snapdragon 8 Elite Gen 6 and agent-native devices — jiqizhixin · 2026-09-23
- Unsloth Desktop Hotfix Adds Qwen-Image-2.1 Image Editing and Fixes GGUF Loading — danielhanchen · 2026-09-23
- Qwen 3.6 35B-A3B Q6 hits ~50 tok/s on a 128GB Strix Halo — what's the best local model now? — jankeydankey · 2026-09-23
- Together AI adds canary rollouts for zero-downtime model upgrades on dedicated inference — togethercompute · 2026-09-23
- Dedicated Hardware for Running AI Agents at Scale Arrives — cyrilzakka · 2026-09-23
- Ternary Bonsai 2 27B: 5.9GB weights retain ~95% of full-precision reasoning — cephaloform · 2026-09-23