Qwen3.8-Flash-Next Runs on Dual DGX Sparks
NVIDIAAI · x · 2026-08-27
MiaAI Lab demonstrated running Qwen3.8-Flash-Next-NVFP4 on two NVIDIA DGX Sparks. The setup supports 900k context with Vision, using SGLang and NVFP4 quantization. Benchmarks show 64 tok/s single stream and 115 tok/s for 2-4 concurrent sessions. An automated script handles weight download, kernel patching, and cluster boot.
More from Infra
- vLLM hits 130k tok/s on DeepSeek V4 Pro in AgentX benchmark — AccBalanced · 2026-08-27
- Google Cloud Run Introduces Instances for MicroVM Deployment — steren · 2026-08-27
- Anthropic Signs $4.5B Compute Deal for Nvidia Rubin Chips at Nscale — Beth_Kindig · 2026-08-27
- DeepSeek-v4-Pro Generates 130M Tokens for $1, 77x Cheaper Than Opus — bookwormengr · 2026-08-27
- NVIDIA FLARE Cuts Federated VLM Training Traffic by 99% — dl_weekly · 2026-08-27
- Same Budget: 256GB Mac or Two DGX Sparks for 70B Inference? — Whyme-__- · 2026-08-27