Running a 320B MoE on two DGX Sparks at 262K context: the 11 things that broke first

TechPreacher · x · 2026-09-29

A hands-on postmortem of serving GLM-5.3-Flash (320B-A18B MoE, NVFP4) at 262K context across two DGX Spark nodes (GB10, 128GB unified memory, 2×200GbE). Key points: 181GiB weights leave 12.9GiB headroom per node for KV and activations, and on GB10 every GPU allocation eats host RAM; stock vLLM simply won't run the model because its NoPE MLA (qkropeheaddim=0) clashes with kernels assuming DeepSeek's pedim=64. It takes a community image with day-0 SM121 fixes (NoPE sparse-MLA via FA2, FlashInfer pinned to 0.6.18 to avoid NaNs, NCCL 2.30.7), images pinned by digest with a verification gate. Eleven fixes in total before it worked.

Original post →

More from Infra

Infra channel →