llama.cpp launch of Qwen3.8 MTP files on 2x DGX Spark cluster hits a dgx-stall — user seeks help
Impossible_Art9151 · reddit · 2026-10-08
A user who successfully runs deepseek-flash on a 2x DGX Spark cluster is stuck launching unsloth's qwen3.8-flash-next MTP version via llama.cpp — the startup command triggers a dgx-stall. The post includes the full launch flags (layer split, draft-dspark speculative decoding, 256k context, RPC) and notes two MTP GGUF files (shared and non-shared) that he can't figure out how to load, asking the community for help.
More from Infra
- MIT Tech Review: building a safer path to autonomous industrial AI — nordicinst · 2026-10-08
- Local inference tuning hits 100+ tok/s; more RAM could push it further — yangyi · 2026-10-08
- Samsung open-sources LittleBit, squeezing a 13B LLM under 1GB with 11.6x inference speedup — ChrSzegedy · 2026-10-08
- Six years on, A100 still leads Nvidia chip mentions in AI papers — ahead of H100+H200 — nathanbenaich · 2026-10-08
- Cloud backlog hits $1.69T as CoreWeave logs $2.58B quarterly revenue — nathanbenaich · 2026-10-08
- Chrome's new DecisionModel API reverse-engineered: prompts, limits and engine tests — dejanseo · 2026-10-08