Red Hat AI ships NVFP4 quantized Qwen3.8-Flash-Next: MoE experts in FP4, vLLM-ready
huggingface · x · 2026-10-05
Red Hat AI published an NVFP4 checkpoint of Qwen3.8-Flash-Next on Hugging Face:
- MoE experts are quantized to FP4 while the rest stays in BF16
- The model is multimodal, accepting text, image, and video inputs
- It's ready to run on vLLM out of the box
- The team also benchmarked it against other popular checkpoints
A ready-made option for cutting VRAM and boosting throughput when self-hosting Qwen on NVIDIA hardware.
More from Infra
- Musk confirms TSMC in talks to build chips at his planned Texas Terafab — Polymarket · 2026-10-06
- Running MiniMax H3 locally on a 12GB GPU: 0.6MP at 4:3 is the sweet spot — vortis23 · 2026-10-06
- AI Data Center Boom Hits Grid Limits as Planned Projects Get Pulled Back — DavidLinthicum · 2026-10-05
- Kirin 9050 reportedly uses 1.5-micron HBI die-to-wafer packaging; analyst says scaling it is the hard part — teortaxesTex · 2026-10-05
- PGlite hits 20 million downloads per week as in-browser Postgres explodes — matei_zaharia · 2026-10-05
- Undergrad's improved GPTQ beats official AWQ on 4-bit Qwen2.5 perplexity — Status-Adeptness8123 · 2026-10-05