Running 2.8T-param Kimi K3 on a 16x GB10 Cluster: 30 tok/s Coding Decode
ciprianveg · reddit · 2026-09-21
A developer shares a hands-on deployment of Moonshot AI's Kimi K3 (2.8T params) on a self-built 16x GB10 cluster:
- Decode: sustained 30 tok/s (peaking 38) during heavy code generation and agentic tasks
- Prefill: 750–910 tok/s, optimized via a modified NCCL topology and dual-switch layout
- Concurrency: multiple concurrent users without dropping generation rates or starving KV cache
- Context: stable multi-hundred-thousand-token runs with repeated 500k compaction
- Network: dual MikroTik CRS804-4DDQ switches with 4x 400G-to-4x100G breakouts
- Runtime: customized gb10-vllm stack with dspark / Inferact Kimi-K3-DSPark wrappers and custom MLA/KV kernels
All runtime patches, configs and build scripts are open-sourced on GitHub (gb10-vllm).
More from Infra
- Tobi Lütke: local Dell server runs DeepSeek 4.1 Flash at ~300 tok/s, a billion tokens a month — BLUECOW009 · 2026-09-21
- Running Qwen3.8-27B EXL3 on RTX 3060 + 5060 Ti: 50 tok/s with tensor parallelism and MTP — bring_back_the_v10s · 2026-09-21
- Baseten CEO says token volume grew 40x YoY while revenue grew ~10x in 12 months — rohanpaul_ai · 2026-09-21
- AI doesn't live in the cloud: who pays the environmental price of scale? — SuzannahB1001 · 2026-09-21
- AI Cluster Bottleneck Isn't Chips — It's the Lasers Moving Data Between Them — McDonaghMatthew · 2026-09-21
- Inside SemiAnalysis: the research firm guiding a $1T AI infrastructure buildout — AccBalanced · 2026-09-21