vLLM hits 464 tok/s on Kimi-K3 with DSpark at batch size 1
vllm_project · x · 2026-07-29
vLLM says it reached a new 464 tok/s peak on Kimi-K3 at batch size 1 under a low-entropy reasoning workload.
- The setup uses 4× GB300 and the public vllm/vllm-openai:kimi-k3 image.
- The benchmark is reproducible with Inferact's DSpark draft model linked in the thread.
- This is mainly a serving-stack / speculative-decoding result for Kimi-K3 on vLLM, not a new model release.
Related event: vLLM Hits 464 tok/s on Kimi K3 with 4 GB300 Systems(2 posts)→
More from Infra
- Google posts first quarter of negative free cash flow after years of growth — michalmalewicz · 2026-07-29
- South Korea’s AI-linked stocks sink 10.84% as Samsung and SK Hynix plunge — emmanuelvivier · 2026-07-29
- OpenAI launches Presence for real-time voice agents and enterprise chatbots — emmanuelvivier · 2026-07-29
- French Startup ZML Releases Free Inference Server Compatible Across AI Chips — emmanuelvivier · 2026-07-29
- AI Chip Startup SambaNova Raises $1B at $11B Valuation — emmanuelvivier · 2026-07-29
- Microsoft is building internal AI models to cut OpenAI dependence and costs by up to 89% — emmanuelvivier · 2026-07-29