Community squeezes a 124B model onto a 128GB DGX Spark with quantization and kernel fixes

alifcoder · x · 2026-09-16

A synthesis of 70 posts from NVIDIA's dev forum on running Ling 3.0 Flash on a single 128GB DGX Spark: community-tested int4/fp4 quantization, MTP, custom kernels, KV cache tuning, SGLang and vLLM benchmarks, plus real failures like memory leaks and their fixes. A showcase of the open-source flywheel turning a doubtful setup into strong performance.

Original post →

More from Infra

Infra channel →