Qwen3.8 Flash hits 415 tok/s on dual DGX Sparks

NVIDIAAI · x · 2026-09-01

Benchmarked Qwen3.8 Flash Next on 2 DGX Sparks using NVFP4 format with 64 concurrent users. It generated 32,768 output tokens in 78.82 seconds, achieving a throughput of 415.7 tok/s. The test passed a 64K/user usable-context stress test with zero leakage, noted as a limit test rather than a production recipe.

Original post →

More from Infra

Infra channel →