Qwen3.8-27B now supports NVFP4 and DFlash2 quantization in SGLang
Alibaba_Qwen · x · 2026-08-21
SGLang cookbook adds NVFP4 + DFlash2 recipes for Qwen3.8-27B. This combination, showing strong community results, offers high-performance W4A4 quantized inference. It supports single-GPU deployment on H200, RTX PRO 6000, and RTX 5090, with detailed installation and configuration guides provided.
More from Infra
- Discussion: Data Privacy in AI Agents and the Case for Private Inference — Many_Audience7660 · 2026-08-21
- Replicating Anthropic Requires 16k Prompt & $250k Synthesis Cost — Ghost_Pilot_MD · 2026-08-21
- Langship: Open Source Tool for Deploying AI Agents Like Terraform — Many_Audience7660 · 2026-08-21
- Ask HN: How Does Qwen 27B Quantized Run on Dual P100 GPUs? — kirisoraa · 2026-08-21
- Pretraining a Mini Kimi K3 on One H200 for $252: A Complete Worklog — joecole · 2026-08-21
- 130 years, 10^22x more compute per dollar: Kurzweil's graph sparks debate — Singularitarian · 2026-08-21