vLLM Hits 464 tok/s on Kimi K3 with 4 GB300 Systems
vLLM announced a new decode throughput peak of 464 tok/s for Kimi K3. Achieved using DSpark on 4 GB300 servers in a low-entropy workload with a batch size of 1, the benchmark highlights significant inference performance gains.
2026-07-29 ~ 2026-07-29 · 2 related posts
- vLLM hits 464 tok/s on Kimi-K3 with DSpark at batch size 1 — vllm_project · 2026-07-29
- vLLM says Kimi K3 hits 464 tok/s decode throughput with DSpark on 4 GB300s — zhyncs42 · 2026-07-29