Running Qwen3.8 Flash NVFP4 on a single DGX Spark: 1M context at 37 tok/s

BLUECOW009 · x · 2026-09-04

Mia AI Lab published a self-contained recipe for serving the 99 GB Qwen3.8-Flash-Next NVFP4 checkpoint on a single DGX Spark (121 GiB unified memory, TP=1) via vLLM with the PLE table offloaded and memory-mapped.

Measured performance:

The author considers it the best model to run on one DGX Spark, beating Qwen3.8-27B and the single-Spark DeepSeek v4 Flash. The repo ships start/stop scripts (10-12 min boot) and documents pitfalls like /dev/shm segment leaks.

Original post →

More from Infra

Infra channel →