180B Model Squeezes onto a Single DGX Spark via 2.39-bit Quantization

A developer released an ultra-low-bit quantization (2.39-bit, experts at 2.125-bit) of the 180B Qwen3.8-Flash model, fitting it on a single DGX Spark with only about 4.5% quality loss and no degradation on long contexts.

2026-10-02 ~ 2026-10-02 · 2 related posts