NVIDIA releases NVFP4-quantized Qwen3.8-Flash-Next: 125B MoE, 63% smaller

_akhaliq · x · 2026-09-05

NVIDIA released the NVFP4-quantized Qwen3.8-Flash-Next on Hugging Face: a 125B MoE with hybrid attention that is now 63% smaller with minimal accuracy loss.

Original post →

More from Models

Models channel →