Running a 320B MoE on two DGX Sparks at 262K context: the 11 things that broke first
TechPreacher · x · 2026-09-29
A hands-on postmortem of serving GLM-5.3-Flash (320B-A18B MoE, NVFP4) at 262K context across two DGX Spark nodes (GB10, 128GB unified memory, 2×200GbE). Key points: 181GiB weights leave 12.9GiB headroom per node for KV and activations, and on GB10 every GPU allocation eats host RAM; stock vLLM simply won't run the model because its NoPE MLA (qkropeheaddim=0) clashes with kernels assuming DeepSeek's pedim=64. It takes a community image with day-0 SM121 fixes (NoPE sparse-MLA via FA2, FlashInfer pinned to 0.6.18 to avoid NaNs, NCCL 2.30.7), images pinned by digest with a verification gate. Eleven fixes in total before it worked.
More from Infra
- Merge Gateway adds self-hosted model support for unified enterprise AI governance — shensi · 2026-09-29
- Chinese CNC factory declines part over export controls; poster builds it anyway — wavefnx · 2026-09-29
- Ben Lorica: the AI data problem has moved downstream — from raw material to usable-data labor and licensed access — bigdata · 2026-09-29
- Researchers Build Backdoor Method to Rank AI Model Energy Use as Big Three Stay Silent — relianceschool · 2026-09-29
- Agent spins up 322 Hugging Face Jobs in 90 minutes for about $4 — victormustar · 2026-09-29
- 80,000+ proxy servers hide stolen AI credentials fueling LLM abuse, Team Cymru finds — ChuckDBrooks · 2026-09-29