vLLM and NVIDIA Achieve Over 25K TPS/GPU for Qwen3.5 on GB200 Systems

AccBalanced · x · 2026-08-08

The vLLM team announced that through deep optimization in collaboration with NVIDIA, they achieved over 25,000 total tokens/s per GPU running Qwen3.5 on the GB200 NVL72 system.

The blog post details the optimization challenges for Qwen3.5's hybrid attention architecture (combining full-attention layers with Gated Delta Network layers). Key technical breakthroughs include:

Original post →

More from Infra

Infra channel →