vLLM says Kimi K3 hits 464 tok/s decode throughput with DSpark on 4 GB300s
zhyncs42 · x · 2026-07-29
- vLLM reports a new peak bs=1 decode throughput on Kimi-K3: 464 tok/s.
- The result comes from a low-entropy reasoning workload running Kimi-K3 + DSpark on 4×4 GB300.
- The team says the benchmark is fully reproducible with the public image vllm/vllm-openai:kimi-k3 and the DSpark draft model linked in the thread.
- The accompanying blog screenshot frames this as a production deployment story: vLLM says Kimi K3 is ready for fast single-user latency, efficient concurrency, and agent-scale serving, with speculative decoding trained via TorchSpec.
Related event: vLLM Hits 464 tok/s on Kimi K3 with 4 GB300 Systems(2 posts)→
More from coding & agent
- Hermes Agent Desktop impresses users with parallel tools and remote local-model setup — Teknium · 2026-07-29
- GrokTerm brings Grok Voice, MCP-ready tabs and audio fixes to Linux — Daniel_Farinax · 2026-07-29
- Using iMessage to text Codex from a plane would keep cloud agents alive offline — rudrank · 2026-07-29
- Slack bot using ChatGPT image generation gets the account banned after 12 hours — jbrov19 · 2026-07-29
- How do you stop GPT from generating junk unit tests all the time? — zeeg · 2026-07-29
- Reddit user looks for a more reliable agent to download and organize IMF reports — Joejoe10x · 2026-07-29