Qwen3.8-27B-FP8 on GH200: 10 concurrent streaming requests, first token in 10ms
MaziyarPanahi · x · 2026-08-15
Lambda user MaziyarPanahi benchmarks Qwen3.8-27B-FP8 on a single NVIDIA GH200 with vLLM: 10 real concurrent streaming requests, each with 16K max output and 262K context. First streamed events arrived within 10ms, and all requests completed normally.
More from Infra
- Next-Gen AI Infra: Beyond GPUs for Energy Efficiency — prateekj · 2026-08-15
- Texas Tightens AI Data Center Approvals: Audits Required for Power, Water, and Community Impact — rohanpaul_ai · 2026-08-15
- Qwen3.8-27B crawls at 5 tokens/s on 8GB VRAM + 32GB RAM: best config? — SoAp9035 · 2026-08-15
- Harvey Trains Custom Model to Cut Costs and Boost Quality in Legal Review — ypatil125 · 2026-08-15
- TrendForce Raises AI Accelerator Shipment Forecast to 31% YoY Growth — Beth_Kindig · 2026-08-15
- NInfer Adds Day-0 Support for Qwen3.8-27B, Hits ~200 tok/s on RTX 5090 — FormOne2615 · 2026-08-15