Full Kimi K3 Model Runs on 16x GB10 Cluster at 38tps Peak

NVIDIAAI · x · 2026-08-06

A developer has successfully run the full Kimi K3 model on a 16x GB10 cluster. Benchmarks show an average generation speed of over 20 tokens/s, a peak of 38 tokens/s, and a prefill speed of 750 tokens/s.

This marks the first successful run of the full K3 model with dspark on this hardware. The developer plans to conduct further optimizations and will release the vllm image and deployment instructions once ready.

Related event: 16x GB10 Cluster Successfully Runs Kimi K3 at 38 TPS Peak(2 posts)→

Original post →

More from Infra

Infra channel →