Dev releases low-bit Qwen3.8-Flash quant keeping 95% bf16 accuracy at long contexts
Crampappydime · reddit · 2026-10-02
A Reddit user released a low-bit quantization of the Qwen3.8-Flash base model on Hugging Face (D JLougen/Qwen3.8-Flash-Next-Mooney).
- Claims 95% of bf16 accuracy retained, with no degradation at longer contexts
- Currently reaches 40 tok/s, with further inference optimizations in progress
- The quant leaves enough headroom to run both 256k context and RoPE
Related event: 180B Model Squeezes onto a Single DGX Spark via 2.39-bit Quantization(2 posts)→
More from Infra
- The Neocloud Reality Check: Why Your Next AI Project May Skip the Big Clouds Entirely — DavidLinthicum · 2026-10-02
- A 3-step guide to open models: picking, hosting locally or via OpenRouter — every · 2026-10-02
- Community inference contest hits 1900 tps prefill, 2.5x faster than omlx baseline — HankYeomans · 2026-10-02
- San Antonio district hosts a dozen data centers as industry camouflages them in woods — The Verge AI · 2026-10-02
- Need a GPU fast? Self-serve clouds like RunPod offer instant spin-ups without contracts — DavidLinthicum · 2026-10-02
- Amazon writes 3,000-word blog warning communities not to block data centers — The Verge AI · 2026-10-02