Qwen3.6 NVFP4 Quantization Speeds Up
danielhanchen · reddit · 2026-07-10
Unsloth released an NVFP4 quantization scheme for Qwen3.6, claiming that without losing precision, it boosted inference speeds for the 27B model to 2.5x that of NVIDIA NVFP4, and the 35B-A3B to 1.56x–1.79x.
The post also provides several benchmark comparisons, including MMLU-Pro, GPQA, AIME 2025, and results against BF16 / FP8 / NVIDIA NVFP4. The author mentioned adding FP8 KV Cache calibration, which automatically doubles the context length to 2x. The text distinguishes between two versions of the 35B model: one optimized for speed (NVFP4-Fast), and another making a slight compromise between speed and precision.
The relevant quantized models are available on Hugging Face, with a more comprehensive analysis and benchmarks in their blog post.
Related event: Unsloth Brings Faster NVFP4 Quantization to Qwen3.6(3 posts)→
More from Infra
- Spomin: live KV cache compaction squeezes 500k tokens of context into 180k resident — wgaca2 · 2026-09-11
- PiPNN nearest-neighbor search wins three awards, up to 78x faster index building — khademinori · 2026-09-11
- M.2-Oculink eGPU Link Silently Downgrades to PCIe Gen1 — Here's How to Check — El_90 · 2026-09-11
- DeepSeek launches V4.1-Flash with 1M-token context and 4x smaller KV-cache — matlabulous · 2026-09-11
- What Can You Still Run on 8GB VRAM? User Asks for Small Models With Tool Use — riceinmybelly · 2026-09-11
- Spain's hourly 80% renewable matching rules clash as France fast-tracks 700MW sites, UK cuts grid queues — eherrerosj · 2026-09-11