Qwen3.8 Flash hits 74 tok/s single-stream, 212 tok/s aggregate on one DGX Spark — open vLLM recipe

DimeRhyme · reddit · 2026-09-29

Results

The author spent weeks tuning a vLLM serving recipe for Qwen3.8 Flash on a single DGX Spark (GB10), fully open-sourced with raw benchmark data:

What made the difference

Footprint

71 GiB model, 16 GB KV pool, 16 GiB free under load; GPU at 35-37W median while generating. Most tricks apply to any MTP/speculative-decoding setup.

Original post →

More from Infra

Infra channel →