Qwen3.8-27B-FP8 on GH200: 10 concurrent streaming requests, first token in 10ms

MaziyarPanahi · x · 2026-08-15

Lambda user MaziyarPanahi benchmarks Qwen3.8-27B-FP8 on a single NVIDIA GH200 with vLLM: 10 real concurrent streaming requests, each with 16K max output and 262K context. First streamed events arrived within 10ms, and all requests completed normally.

Original post →

More from Infra

Infra channel →