llama.cpp launch of Qwen3.8 MTP files on 2x DGX Spark cluster hits a dgx-stall — user seeks help

Impossible_Art9151 · reddit · 2026-10-08

A user who successfully runs deepseek-flash on a 2x DGX Spark cluster is stuck launching unsloth's qwen3.8-flash-next MTP version via llama.cpp — the startup command triggers a dgx-stall. The post includes the full launch flags (layer split, draft-dspark speculative decoding, 256k context, RPC) and notes two MTP GGUF files (shared and non-shared) that he can't figure out how to load, asking the community for help.

Original post →

More from Infra

Infra channel →