Reddit user shows NVFP4 Qwen3.8-27B matches BF16 with two sampling tweaks, runs 3x faster

UmpireBorn3719 · reddit · 2026-09-07

A Reddit user benchmarked Qwen3.8-27B locally on an RTX 5090, comparing NInfer+NVFP4 (4-bit quant, Unsloth Dynamic weights) vs llama.cpp+Q5KM on IFBench (50 samples).

Key findings:

Takeaway: NVFP4 failures are mostly sampling noise; with two tweaked params you get near-FP16 instruction following at 3x speed and far less VRAM. The strict/loose gap is a useful diagnostic for noise vs real failures. Author notes n=50 variance but consistent direction across reruns.

Original post →

More from Infra

Infra channel →