Red Hat AI ships NVFP4 quantized Qwen3.8-Flash-Next: MoE experts in FP4, vLLM-ready

huggingface · x · 2026-10-05

Red Hat AI published an NVFP4 checkpoint of Qwen3.8-Flash-Next on Hugging Face:

A ready-made option for cutting VRAM and boosting throughput when self-hosting Qwen on NVIDIA hardware.

Original post →

More from Infra

Infra channel →