NVFP4 quantization matches bf16 and FP8 in NVIDIA's own evals, Redditor argues
General-Spite1222 · reddit · 2026-10-09
A Redditor makes the case that NVIDIA's NVFP4 quantization format is on par with Q8, citing NVIDIA's own evaluations:
- Official evals show Qwen3.8-27B-NVFP4 matching bf16 and Qwen3.8-Flash-Next-NVFP4 matching FP8
- The author notes plenty of community trash talk about NVFP4 but no numbers backing it up
The thread reignites the quantization-quality debate: if NVFP4 really is near-bf16, it implies major VRAM and inference cost savings for local deployment and serving stacks.
More from Infra
- Atomic Agent Desktop: free MIT-licensed open-source local AI agent app — SucceededMind · 2026-10-09
- Samsung expects 780% quarterly operating profit jump on AI boom — chemist_slime · 2026-10-09
- TRL v1.15 ships fused LM head: 82% less peak VRAM, 7x longer sequences — QGallouedec · 2026-10-09
- Dev streams a 66GB unquantized 22B video model on an iGPU with only 15.6GB shared RAM — Business_Swordfish_5 · 2026-10-09
- Asia-Pacific Needs $280 Billion to Build 26.5 GW of Planned Data Centers, Johor Leads — shashib · 2026-10-09
- Stepped MoE paper: one model scales from 1B to 4B parameters, beating dense counterparts by 2-5% — pmttyji · 2026-10-09