Qwen3.8-Flash recipe on DGX Spark: int4 quantization + RDMA, 47.5 t/s code generation

Saren-WTAKO · reddit · 2026-08-31

The author shares a recipe for running Qwen3.8-Flash on a single GB10/DGX Spark, using Intel AutoRound int4 quantization, vLLM, and offloading fp8 ngram table to local SSD or external RDMA server. Detailed performance numbers: 47.5 t/s for code, 60 t/s for JSON, and discusses MTP-induced TTFT overhead.

Original post →

More from Infra

Infra channel →