Fastest NVFP4 quant of Qwen3.8 27B released, 50% faster than Q4 on compatible hardware
ionsago · reddit · 2026-08-21
A new Blackwell-native, prefill-optimized NVFP4 4-bit quantization of Qwen3.8 27B has been released. It runs 50% faster than a standard Q4 quant of the same memory footprint. Benchmarked on RTX 5090 32GB, it achieves 6250 t/s, outperforming other NVFP4 quants by 4-7%. The GGUF also includes a quantized MTP draft head, offering an additional 15% speed boost with recommended settings.
More from Models
- API Model Gains Vision Capabilities in Major Update — teortaxesTex · 2026-08-21
- Experts question Anthropic's trust in Claude's simulation excuses — GarrisonLovely · 2026-08-21
- LLM German Output Cringed: Reads Like It Was Written by Olaf Scholz — DominiqueCAPaul · 2026-08-21
- DeepSeek V4 Flash adds vision capabilities, now live — TheCryptoCat75 · 2026-08-21
- DeepSeek V4-vision-exp launches API with ultra-fast speed and low cost — teortaxesTex · 2026-08-21
- Jie Tang on scaling history: FLOPs were intelligence, parameters were knowledge — cedric_chee · 2026-08-21