QUASAR Releases Fully Quantized NVFP4 Qwen3.8-27B, Near-BF16 Performance
arty_photography · reddit · 2026-08-26
QUASAR team releases a fully quantized NVFP4 version of Qwen3.8-27B, trained with quantization-aware distillation. The model size drops from 55.6GB to 19.7GB while maintaining near-BF16 performance on GPQA-Diamond and AIME26, outperforming other NVFP4 quantizations. Supports vLLM on Blackwell GPUs. Paper available.
More from Models
- Nvidia may have funded 100T tokens for free GLM 5.3 Flash release — bindureddy · 2026-08-26
- Flaw in anti-finetuning: Cost > Quality once models are saturated — rhythmrg · 2026-08-26
- sanoTTS: 1.4M-Param Model Runs Real-Time on $3 Chip — kastnerkyle · 2026-08-26
- New Models to Know: MoE-ViE, τ0-VLA, 4DAnyone, and More — TheTuringPost · 2026-08-26
- User Finds Sol Max More Reliable Than Sol Ultra for Complex Tasks — imjustnewatai · 2026-08-26
- GPT Auto-Titles Conversation in Chinese, Baffling User — rodrigoinfloripa · 2026-08-26