Qwen3.8-27B hits 129.8 tok/s on one GH200; vLLM wins

MaziyarPanahi · x · 2026-08-21

Benchmarks show Qwen3.8-27B BF16 achieving 129.8 tok/s on a single NVIDIA GH200. The setup included 2,048 input tokens and 1,536 forced output tokens. For single-stream requests, vLLM + DFlash2 was the clear winner, offering 2.18x speedup over plain autoregressive and 9.4% improvement over MTP-3. Compute credits provided by LambdaAPI.

Related event: Qwen3.8-27B Hits 129.8 tok/s on a Single GH200(2 posts)→

Original post →

More from Infra

Infra channel →