Qwen3.8-27B hits 129.8 tok/s on a single NVIDIA GH200 with vLLM
MaziyarPanahi · x · 2026-08-21
Qwen3.8-27B (BF16) achieved 129.8 tokens/s on a single NVIDIA GH200. The setup utilizes vLLM and DFlash2, showing a 2.18x speedup over plain autoregressive inference and a 9.4% improvement over MTP-3 for single-stream requests.
Related event: Qwen3.8-27B Hits 129.8 tok/s on a Single GH200(3 posts)→
More from Infra
- DecagonAI achieves sub-30ms latency for real-time TTS — dhruv2038 · 2026-08-21
- US union warns: New England data center bans could kill thousands of jobs — Polymarket · 2026-08-21
- Waymo Cuts Hardware to $20k, Tesla Aims for Cybercab COGS Under $20k — JOBhakdi · 2026-08-21
- Cursor's Git Storage System: S3 as Source of Truth, Local Disk as Cache — xennygrimmato_ · 2026-08-21
- Unsloth Desktop Update: Auto Compaction and LAN Remote Access — danielhanchen · 2026-08-21
- AMD ROCm 10.1 fixes major issues: LLaMA.cpp runs flawlessly on RDNA2 — smellof · 2026-08-21