QUASAR Releases Fully Quantized NVFP4 Qwen3.8-27B, Near-BF16 Performance

arty_photography · reddit · 2026-08-26

QUASAR team releases a fully quantized NVFP4 version of Qwen3.8-27B, trained with quantization-aware distillation. The model size drops from 55.6GB to 19.7GB while maintaining near-BF16 performance on GPQA-Diamond and AIME26, outperforming other NVFP4 quantizations. Supports vLLM on Blackwell GPUs. Paper available.

Original post →

More from Models

Models channel →