Qwen3.8-27B hits 129.8 tok/s on a single NVIDIA GH200 with vLLM

MaziyarPanahi · x · 2026-08-21

Qwen3.8-27B (BF16) achieved 129.8 tokens/s on a single NVIDIA GH200. The setup utilizes vLLM and DFlash2, showing a 2.18x speedup over plain autoregressive inference and a 9.4% improvement over MTP-3 for single-stream requests.

Related event: Qwen3.8-27B Hits 129.8 tok/s on a Single GH200(3 posts)→

Original post →

More from Infra

Infra channel →