Qwen3.8 hits 181 tok/s aggregate on dual DGX Spark nodes
Developers deployed Qwen3.8-Flash-Next on a 2-node NVIDIA DGX Spark (GB10) with NVFP4 quantization, achieving 181 tok/s aggregate throughput, with about 50 tok/s decode and 2900 tok/s prefill.
2026-08-29 ~ 2026-08-30 · 2 related posts
- Achieving 181 tok/s on Qwen3.8 with 2x DGX Sparks via NVMe offloading — StartupTim · 2026-08-29
- Qwen3.8-Flash-Next on 2x DGX Spark NVFP4: 50 t/s decode, 2,900 t/s prefill — -dysangel- · 2026-08-30