Qwen3.6 NVFP4 Quantization Speeds Up
danielhanchen · reddit · 2026-07-10
Unsloth released an NVFP4 quantization scheme for Qwen3.6, claiming that without losing precision, it boosted inference speeds for the 27B model to 2.5x that of NVIDIA NVFP4, and the 35B-A3B to 1.56x–1.79x.
The post also provides several benchmark comparisons, including MMLU-Pro, GPQA, AIME 2025, and results against BF16 / FP8 / NVIDIA NVFP4. The author mentioned adding FP8 KV Cache calibration, which automatically doubles the context length to 2x. The text distinguishes between two versions of the 35B model: one optimized for speed (NVFP4-Fast), and another making a slight compromise between speed and precision.
The relevant quantized models are available on Hugging Face, with a more comprehensive analysis and benchmarks in their blog post.
Related event: Unsloth Brings Faster NVFP4 Quantization to Qwen3.6(3 posts)→
More from Infra
- Vercel AI Gateway data shows Anthropic, OpenAI and Google at 97.09% spend share — cramforce · 2026-07-21
- NVIDIA starts rolling out 102.4 Tbps Spectrum-6 switches for Vera Rubin AI factories — nvidia · 2026-07-21
- Mustafa Suleyman says Microsoft is preparing for an OpenAI exit, while a new chip costs 30% less than GB200 — thoefler · 2026-07-21
- Microsoft and Mistral sign multi-billion-dollar deal to expand AI infrastructure in Europe — The Decoder · 2026-07-21
- Speculative decoding boosts Qwen3.6-27B on one 5090, but slows crowded servers — luke_pacman · 2026-07-21
- NVIDIA says Blackwell Ultra hit 1,648 TFLOPs per GPU on DeepSeek-V3 671B training — NVIDIAAI · 2026-07-21