Qwen3.8-27B hits 129.8 tok/s on one GH200; vLLM wins
MaziyarPanahi · x · 2026-08-21
Benchmarks show Qwen3.8-27B BF16 achieving 129.8 tok/s on a single NVIDIA GH200. The setup included 2,048 input tokens and 1,536 forced output tokens. For single-stream requests, vLLM + DFlash2 was the clear winner, offering 2.18x speedup over plain autoregressive and 9.4% improvement over MTP-3. Compute credits provided by LambdaAPI.
Related event: Qwen3.8-27B Hits 129.8 tok/s on a Single GH200(2 posts)→
More from Infra
- AI Energy Crisis: Grid, Not Cost, Is the Bottleneck; Solar and Batteries Offer Fastest Path — PeterDiamandis · 2026-08-21
- SmashTable: In-memory containers with DBMS-like transactions open-sourced — srchvrs · 2026-08-21
- Marin Uses Scaling Laws to Predict Training Trajectories, Kicks Off 535B MoE — dlwh · 2026-08-21
- Qwen3.8-27B achieves 2x decode speedup with DFlash2 on 256k context — maddie-lovelace · 2026-08-21
- Miles v0.1 integrates Mooncake backend, accelerating remote data fetch by 14x — BanghuaZ · 2026-08-21
- AlloyDB scales vector search to 10 billion vectors with four-level tree — rseroter · 2026-08-21