180B Qwen model runs on one DGX Spark: 2.39-bit quant keeps 95.5% of BF16 scores

TheZachMueller · x · 2026-10-02

Developer DJLougen released an extreme low-bit quantization of Qwen3.8-Flash-Next, now on Hugging Face.

The upshot: a 180B-parameter MoE model running on a single DGX Spark, with thanks to Lambda and HF's Zach Mueller for compute support.

Original post →

More from Infra

Infra channel →