vLLM 0.28.0 Perf: Kimi-K3 Optimizations & Hardware Adaptation
vllm_project · x · 2026-08-27
vLLM v0.28.0 is released with significant performance improvements and hardware adaptations:
- Kimi-K3: Adaptive speculative token budget improves E2E TTFT by 5565%; optional shared-expert sharding saves 17 GiB per GPU.
- Sequence Parallelism: Combined all-gathers offer 1.53x speedup at the kernel level.
- NVIDIA: FlashInfer XQA decode on SM12x; native DSA path for MTP=3 on SM90.
- AMD ROCm: Supports torch 2.12 / triton 3.7, plus a fused Kimi-K3 KDA decode kernel.
- Intel XPU: Added a torch linear backend with blockwise GEMM; XPU wheel in release pipeline.
- CPU: Added MLA backend to run DeepSeek-V2/V3 on CPU.
Related event: vLLM v0.28.0 Ships with 584 Commits from 270 Contributors(3 posts)→
More from Infra
- Dev forks Nvidia drivers to enable PCIe P2P on GeForce for SlimServe — QuixiAI · 2026-08-27
- View: Single dev with 200 B300s could beat Alibaba's post-training team — kalomaze · 2026-08-27
- GLM 5.3 Flash Benchmark: Hits 881 tok/s on Dual DGX — teortaxesTex · 2026-08-27
- Rakyll advises: If you have any CPU nodes, hold on to them — rakyll · 2026-08-27
- OpenRouter Serves 200B Total Tokens, Adds Qwen 3.5 35B — gajesh · 2026-08-27
- $300 AMD BC-250 runs 35B model at 67 tok/s: Local AI hardware barrier collapses — SimplyAnnisa · 2026-08-27