180B Qwen model runs on one DGX Spark: 2.39-bit quant keeps 95.5% of BF16 scores
TheZachMueller · x · 2026-10-02
Developer DJLougen released an extreme low-bit quantization of Qwen3.8-Flash-Next, now on Hugging Face.
- 2.125-bit experts, 2.39 bits/weight across the transformer
- 92 GB download, 39 GiB in memory
- Retains 95.5% of BF16's benchmark scores
The upshot: a 180B-parameter MoE model running on a single DGX Spark, with thanks to Lambda and HF's Zach Mueller for compute support.
More from Infra
- Cloudflare AI Search goes GA with multimodal embeddings, PDF OCR, billing from Nov 1 — michellechen · 2026-10-02
- TailSlayer: open-source C++ library cuts RAM read tail latency with hedged reads across DRAM channels — lauriewired · 2026-10-02
- DataScalar: The Failed 90s Memory-Centric Architecture That Aged Well — lauriewired · 2026-10-02
- Frontier Airlines to offer Starlink WiFi starting early 2027 across 1,000+ aircraft — elonmusk · 2026-10-02
- Will Cloud AI Agents Need Residential IPs? Amazon Blocked Meta's Muse — SignificantFail3632 · 2026-10-02
- Open Anonymity launches OA-chat and zkAPI, a zero-knowledge protocol co-built with Ethereum Foundation — kenziyuliu · 2026-10-02