Fastest NVFP4 quant of Qwen3.8 27B released, 50% faster than Q4 on compatible hardware

ionsago · reddit · 2026-08-21

A new Blackwell-native, prefill-optimized NVFP4 4-bit quantization of Qwen3.8 27B has been released. It runs 50% faster than a standard Q4 quant of the same memory footprint. Benchmarked on RTX 5090 32GB, it achieves 6250 t/s, outperforming other NVFP4 quants by 4-7%. The GGUF also includes a quantized MTP draft head, offering an additional 15% speed boost with recommended settings.

Original post →

More from Models

Models channel →