Dev releases low-bit Qwen3.8-Flash quant keeping 95% bf16 accuracy at long contexts

Crampappydime · reddit · 2026-10-02

A Reddit user released a low-bit quantization of the Qwen3.8-Flash base model on Hugging Face (D JLougen/Qwen3.8-Flash-Next-Mooney).

Related event: 180B Model Squeezes onto a Single DGX Spark via 2.39-bit Quantization(2 posts)→

Original post →

More from Infra

Infra channel →