Qwen3.8-27B Hits 129.8 tok/s on a Single GH200

Benchmark tests show Qwen3.8-27B (BF16) achieving 129.8 tokens/s on a single NVIDIA GH200 using vLLM with the DFlash2 stack, with vLLM outperforming in the comparison.

2026-08-21 ~ 2026-08-21 · 2 related posts

1 near-duplicate retellings: MaziyarPanahi