Quantized Qwen3.8 draft model cuts VRAM 2.6x, boosts context from 90K to 130K on one RTX 5090
QuixiAI · x · 2026-08-30
- Developer nb4ld's first HF upload: calibrated NVFP4 (W4A4) quantization of the DFlash 2 block-diffusion speculative decoding draft for Qwen3.8-27B, in NVIDIA ModelOpt layout that SGLang loads natively.
- Built for a single RTX 5090 (32GB): the draft shrinks from 3.53GB (BF16) to 1.37GB, turning into +44% KV-cache context (90K → 130K tokens) at the same decode speed (215 → 228 tok/s end-to-end) and same acceptance rate (3.71 → 3.60, within ±0.2 noise).
- Output quality is unchanged by construction since the target model verifies every drafted token; uncalibrated RTN quantization drops acceptance to 3.26, showing calibration matters.
- Also benchmarked a coding-agent session with tool calls on the SGLang source tree, with decode matching BF16 across context sizes.
More from Infra
- 40-nm Memristor Chip Turns Conductance Drift Into a Feature, Beats A100 by 50-480x — maier_ak · 2026-09-01
- 40nm Neural-Dynamics Chip Uses Conductance Drift for 2.12ms Iteration Latency — maier_ak · 2026-09-01
- Qwen3.8 Flash hits 415 tok/s on dual DGX Sparks — NVIDIAAI · 2026-09-01
- OpenAI's 'Jalapeno' Chip Revealed: 1500 Tokens/s Throughput — firstadopter · 2026-09-01
- TensorSharp vs llama.cpp: Qwen 3.8 Flash Next Benchmarks — fuzhongkai · 2026-09-01
- Tencent Hunyuan AngelSlim: Compressing Hy4 Model to 214GB with Heterogeneous Inference — 腾讯混元 · 2026-09-01