vLLM says Kimi K3 hits 464 tok/s decode throughput with DSpark on 4 GB300s
zhyncs42 · x · 2026-07-29
- vLLM reports a new peak bs=1 decode throughput on Kimi-K3: 464 tok/s.
- The result comes from a low-entropy reasoning workload running Kimi-K3 + DSpark on 4×4 GB300.
- The team says the benchmark is fully reproducible with the public image vllm/vllm-openai:kimi-k3 and the DSpark draft model linked in the thread.
- The accompanying blog screenshot frames this as a production deployment story: vLLM says Kimi K3 is ready for fast single-user latency, efficient concurrency, and agent-scale serving, with speculative decoding trained via TorchSpec.
Related event: vLLM Hits 464 tok/s on Kimi K3 with 4 GB300 Systems(2 posts)→
More from coding & agent
- A Gemini agent to auto-reset your 50+ leaked passwords: a killer use case — sup_nim · 2026-09-23
- OpenAI startup engineering lead: in 2026 'everything is a coding agent' — simple and elegant wins — RichmanRonald · 2026-09-23
- Dev building Infinite Craft clone on Roblox finds Gemini Flash terrible, asks for model picks — DisastrousUpstairs23 · 2026-09-23
- This setup keeps a spare iPhone on the desk so one agent can drive both Mac and phone — signulll · 2026-09-23
- Agent design rule: verifiers may give feedback but never promote candidates — blaizedsouza · 2026-09-23
- AI engineering is more like lawmaking than board games, argues Drew Breunig — dbreunig · 2026-09-23