Full Kimi K3 Model Runs on 16x GB10 Cluster at 38tps Peak
NVIDIAAI · x · 2026-08-06
A developer has successfully run the full Kimi K3 model on a 16x GB10 cluster. Benchmarks show an average generation speed of over 20 tokens/s, a peak of 38 tokens/s, and a prefill speed of 750 tokens/s.
This marks the first successful run of the full K3 model with dspark on this hardware. The developer plans to conduct further optimizations and will release the vllm image and deployment instructions once ready.
Related event: 16x GB10 Cluster Successfully Runs Kimi K3 at 38 TPS Peak(2 posts)→
More from Infra
- Startup Panthalassa Develops Floating AI Data Centers Powered by Wave Energy — TinfoilTricorn · 2026-08-06
- Maple 20B Hits 9.8k Tokens/s on Single GH200, Opens Free Inference Endpoint — MaziyarPanahi · 2026-08-06
- Maple 20B Hits 9,885 tokens/s on Single GH200 with 64 Concurrent Requests — MaziyarPanahi · 2026-08-06
- Open-Source LLM Inference Engine TokenSpeed Joins PyTorch Ecosystem — zhyncs42 · 2026-08-06
- Stranded Gas Wells to Power AI: Converting Methane into Methanol for Data Centers — Ghost_Pilot_MD · 2026-08-06
- FPTalks 2026 Focuses on Low-Precision LLM Pretraining — nmwsharp · 2026-08-06