vLLM hits 464 tok/s on Kimi-K3 with DSpark at batch size 1
vllm_project · x · 2026-07-29
vLLM says it reached a new 464 tok/s peak on Kimi-K3 at batch size 1 under a low-entropy reasoning workload.
- The setup uses 4× GB300 and the public vllm/vllm-openai:kimi-k3 image.
- The benchmark is reproducible with Inferact's DSpark draft model linked in the thread.
- This is mainly a serving-stack / speculative-decoding result for Kimi-K3 on vLLM, not a new model release.
Related event: vLLM Hits 464 tok/s on Kimi K3 with 4 GB300 Systems(2 posts)→
More from Infra
- Qualcomm goes agent-centric: Snapdragon 8 Elite Gen 6 and agent-native devices — jiqizhixin · 2026-09-23
- Unsloth Desktop Hotfix Adds Qwen-Image-2.1 Image Editing and Fixes GGUF Loading — danielhanchen · 2026-09-23
- Qwen 3.6 35B-A3B Q6 hits ~50 tok/s on a 128GB Strix Halo — what's the best local model now? — jankeydankey · 2026-09-23
- Together AI adds canary rollouts for zero-downtime model upgrades on dedicated inference — togethercompute · 2026-09-23
- Dedicated Hardware for Running AI Agents at Scale Arrives — cyrilzakka · 2026-09-23
- Ternary Bonsai 2 27B: 5.9GB weights retain ~95% of full-precision reasoning — cephaloform · 2026-09-23