NVIDIA Releases NVFP4 Quantized Qwen3.8-Flash-Next
NVIDIA published an NVFP4-quantized version of Qwen3.8-Flash-Next on Hugging Face. The 125B-parameter MoE model with hybrid attention supports text and image input, with its size reduced by 63% after quantization.
2026-09-05 ~ 2026-09-05 · 2 related posts
- NVIDIA releases NVFP4-quantized Qwen3.8-Flash-Next: 125B MoE, 63% smaller — _akhaliq · 2026-09-05
- NVIDIA's NVFP4-quantized Qwen3.8-Flash-Next image-text-to-text model trends on Hugging Face — nvidia · 2026-09-05