NVIDIA Releases NVFP4 Quantized Qwen3.8-Flash-Next

NVIDIA published an NVFP4-quantized version of Qwen3.8-Flash-Next on Hugging Face. The 125B-parameter MoE model with hybrid attention supports text and image input, with its size reduced by 63% after quantization.

2026-09-05 ~ 2026-09-05 · 2 related posts