vLLM Hits 464 tok/s on Kimi K3 with 4 GB300 Systems

vLLM announced a new decode throughput peak of 464 tok/s for Kimi K3. Achieved using DSpark on 4 GB300 servers in a low-entropy workload with a batch size of 1, the benchmark highlights significant inference performance gains.

2026-07-29 ~ 2026-07-29 · 2 related posts