vLLM and NVIDIA Achieve Over 25K TPS/GPU for Qwen3.5 on GB200 Systems
AccBalanced · x · 2026-08-08
The vLLM team announced that through deep optimization in collaboration with NVIDIA, they achieved over 25,000 total tokens/s per GPU running Qwen3.5 on the GB200 NVL72 system.
The blog post details the optimization challenges for Qwen3.5's hybrid attention architecture (combining full-attention layers with Gated Delta Network layers). Key technical breakthroughs include:
- Blackwell-Optimized GDN Prefill Kernel: Delivering a 1.02x to 5.78x performance improvement over previous FLA/Triton implementations across various model sizes and batch shapes.
- Disaggregated Serving State Transfer: Solving the challenge of correctly transferring heterogeneous attention and GDN states between prefill and decode workers.
More from Infra
- AI Data Center Firm Switch Confidentially Files for US IPO — Polymarket · 2026-08-08
- MIT's 'Implosion Carving' Shrinks 3D Photonic Devices to Channel Visible Light — snikolov · 2026-08-08
- Nvidia B300 Specs Contradiction: More SMs but Same BF16 TFLOPS as B200 — StasBekman · 2026-08-08
- Model Routing Reshapes AI Economics: Glean Cuts Latency 50% and Speeds Search 10x — VibeMarketer_ · 2026-08-08
- Running Qwen 3.6 27B on RTX 5090: 40 t/s at 262k Context in llama.cpp — Gargle-Loaf-Spunk · 2026-08-08
- Red Hat Releases New DSpark Models, Boosting vLLM Inference Speed by 4x — vllm_project · 2026-08-08