Qwen3.8-Flash-Next Runs on Dual DGX Sparks

NVIDIAAI · x · 2026-08-27

MiaAI Lab demonstrated running Qwen3.8-Flash-Next-NVFP4 on two NVIDIA DGX Sparks. The setup supports 900k context with Vision, using SGLang and NVFP4 quantization. Benchmarks show 64 tok/s single stream and 115 tok/s for 2-4 concurrent sessions. An automated script handles weight download, kernel patching, and cluster boot.

Original post →

More from Infra

Infra channel →