DGX Spark struggles with dense models; MoE better suited for unified memory

TheZachMueller · x · 2026-08-09

A developer reports that dense models like Qwen 3.8 27B perform terribly on unified-memory systems like DGX Spark, while MoE models with few active parameters per token are better suited. A 200K-context agentic session on a 200B-parameter model with 10B active parameters would take 22 minutes on 2x DGX Sparks vs 2-3 minutes on 2x RTX PRO 6000s, with worse scaling for concurrency and tensor parallelism. Conclusion: GPUs outperform unified memory for anything beyond a chat interface.

Related event: Dual RTX PRO 6000 Outperforms DGX Spark by 10x in Local LLM Tests(2 posts)→

Original post →

More from Infra

Infra channel →