Qwen3.8-Flash-Next on 2x DGX Spark NVFP4: 50 t/s decode, 2,900 t/s prefill

-dysangel- · reddit · 2026-08-30

Redditor -dysangel- shares a full config for running Qwen3.8-Flash-Next NVFP4 on 2x NVIDIA DGX Spark (GB10): 49.7 t/s decode on structured output / 34.8 t/s prose, 2,875 t/s prefill on 11k tokens under TP2; single node does 35 t/s.

Stack highlights:

Dead ends documented:

Next steps: block-diffusion drafting (DFlash/DSpark) and disaggregated decode onto a big-bandwidth Mac. MTP acceptance is 3.8 tokens/step on structured output but noticeably lower on prose — benchmark both.

Related event: Qwen3.8 hits 181 tok/s aggregate on dual DGX Spark nodes(2 posts)→

Original post →

More from Infra

Infra channel →