Qwen3.8-27B NVFP4 quants compared: lm_head precision makes or breaks MTP speedups
danielhanchen · x · 2026-09-14
- Daniel Han of Unsloth compared multiple NVFP4 quantizations of Qwen3.8-27B: accuracy is close across variants; the real differences are memory and speed.
- For VRAM-limited setups, minima-ai/mnmaqwen3.827bnvfp4 is a good pick but lacks MTP. NVIDIA's variant includes MTP, but in long-context coding tests MTP-4 is only 2x faster than no MTP and still 2.5–3x slower than Unsloth on an RTX Pro 6000.
- Likely culprit: NVIDIA quantizes lmhead to NVFP4 while Unsloth keeps it FP8. Since MTP shares the target model's lmhead, this can hurt prediction quality and acceptance rate.
- Verdict: Unsloth NVFP4 is the pick; long-horizon agentic coding accuracy tests are ongoing.
More from Infra
- RTX 5090 Listed at $6,000 CAD at Canada Computers — yacineMTB · 2026-09-14
- Got a 384GB RAM, dual RTX 6000 workstation for $2400 — now what? — ABDULLAH3_33 · 2026-09-14
- Nvidia partners to build 2GW of AI capacity in Australia by 2027, more than doubling 1.6GW — Beth_Kindig · 2026-09-14
- MiniMax 3-step Turbo LoRA tested: 2-min audio+video gen on RTX 5090 — agapes1270 · 2026-09-14
- Musk says SpaceX will launch Nvidia AI computers into space next year — inductionheads · 2026-09-14
- Procedurally Generated FlashAttention With CuTe Tilings Yields Runnable CUDA Code — vtabbott_ · 2026-09-14