16x GB10 Cluster Successfully Runs Kimi K3 at 38 TPS Peak

A developer successfully ran the complete Kimi K3 model on a 16x GB10 cluster. Tests revealed impressive performance, with an average generation speed exceeding 20 tokens/s, a peak of 38 tokens/s, and a prefill speed of 750 tokens/s.

2026-08-05 ~ 2026-08-06 · 2 related posts

1 near-duplicate retellings: NVIDIAAI