Achieving 181 tok/s on Qwen3.8 with 2x DGX Sparks via NVMe offloading

StartupTim · reddit · 2026-08-29

A developer achieved 181 tok/s aggregate throughput on Qwen3.8-Flash-Next using a 2-node NVIDIA DGX Spark (GB10) cluster, with single-stream decode at 30–50 tok/s.

Key Optimizations:

Original post →

More from coding & agent

coding & agent channel →