vLLM Adds Hardware & Performance Optimizations
vllm_project · x · 2026-07-12
Supplementary hardware and performance notes for vLLM 0.25.0, highlighting cross-platform optimizations:
- NVIDIA / GB300: FlashInfer's fused all-reduce tuned for GB300, Blackwell's NVFP4 decode throughput restored, and FA4-MLA warmup & XQA decode kernels added.
- GLM-5.2 / DeepSeek: fused indexer kernel brings 1.9–3.3% end-to-end throughput boost, reduce-scatter MoE all-reduce improves 3%, and DSv4's tokentoreqindices cache accelerates certain kernels by 5–6x.
- AMD ROCm: torch 2.11 stable ABI, AITER FlashAttention MLA prefill backend, and shared-expert fused kernel.
- Intel XPU: W8A8 FP8 linear kernel, multi-granularity quantization, and uniform-batch CUDA graphs for FA2.
- CPU / Other architectures: Faster unquantized MoE on AArch64, Apple Silicon hang fixed, and support for RISC-V RVV INT4 GEMM & PowerPC fp16.
More from Infra
- Production AI budgets include retries, routing, caching and observability—not just token prices — arx-go · 2026-07-22
- NVIDIA briefs analysts on Vera CPU and doubles down on monolithic agentic design — BenBajarin · 2026-07-22
- NVIDIA unveils Vera Rubin platform with claims of 10x better performance per watt — nvidia · 2026-07-22
- SkyPilot comes out of stealth with a pitch to unify fragmented AI compute — skypilot_org · 2026-07-22
- Why a 1GW Chinese AI data center may be plausible after all — teortaxesTex · 2026-07-22
- AI agent accountability layer adds terminal verification with explicit finality and no signup — Special_Librarian145 · 2026-07-22