Reddit user shows NVFP4 Qwen3.8-27B matches BF16 with two sampling tweaks, runs 3x faster
UmpireBorn3719 · reddit · 2026-09-07
A Reddit user benchmarked Qwen3.8-27B locally on an RTX 5090, comparing NInfer+NVFP4 (4-bit quant, Unsloth Dynamic weights) vs llama.cpp+Q5KM on IFBench (50 samples).
Key findings:
- With default sampling (temp=1.0, minp=0), Q5KM scored 76% strict vs NVFP4's 74% — seemingly a Q5 win.
- But NVFP4's 4-point strict/loose gap (74/78) showed its failures were mostly formatting noise (indent, casing, word count), while Q5's zero gap meant genuine capability failures.
- Tightening sampling to temp=0.9, minp=0.05 flipped the results: NVFP4 hit 80% strict/loose, matching local BF16 and approaching the official 300-sample BF16 baseline of 79.5%. Q5 stayed at 76%.
Takeaway: NVFP4 failures are mostly sampling noise; with two tweaked params you get near-FP16 instruction following at 3x speed and far less VRAM. The strict/loose gap is a useful diagnostic for noise vs real failures. Author notes n=50 variance but consistent direction across reruns.
More from Infra
- Custom llama.cpp Branch Adds Expert Expansion for MoE Models — Specific-Tax-6700 · 2026-09-07
- Developer earns just $2.60 per cycle running a bot on OpenAI-subsidized tokens — TheMoonMidas · 2026-09-07
- Interactive Speculative Decoding Tutorial for NeurIPS Explains When It Stays Lossless — Madisonkanna · 2026-09-07
- Dell COO: AI boom means shortages across DRAM, NAND, CPUs, disks and more — mattwbaker · 2026-09-07
- 2× R9700 local rig runs Qwen3.8-27B at 111 tok/s for ~€4k, €1k under a 5090 — smallDeltaBigEffect · 2026-09-07
- Budget GPU Advice for Local LLMs: Modded 2080 Ti vs Mi50 Under $700 — Current-Set1963 · 2026-09-07