16x GB10 Cluster Successfully Runs Kimi K3 at 38 TPS Peak
A developer successfully ran the complete Kimi K3 model on a 16x GB10 cluster. Tests revealed impressive performance, with an average generation speed exceeding 20 tokens/s, a peak of 38 tokens/s, and a prefill speed of 750 tokens/s.
2026-08-05 ~ 2026-08-06 · 2 related posts
- Full Kimi K3 model runs on 16x GB10 cluster at 20+ TPS — ciprianveg · 2026-08-05
1 near-duplicate retellings: NVIDIAAI