Swapping Qwen's N-gram layer to Q8 shows no real speed penalty in local inference tests

Altruistic_Heat_9531 · reddit · 2026-09-03

A local-LLM experimenter replaced the low-precision N-gram portion of an IQ4XS Qwen 3.8 Next model with Q8 weights (following someone else's BF16 experiment) and benchmarked on a Xeon E5-2690v4 + RTX 3090 (250W cap) with 96GB DDR4, no MTP.

Short-run generation jumped from 8.8 to 10.7 tok/s, but both setups converge to a steady state around 10.1 tok/s — meaning no meaningful speed penalty from the higher-precision N-gram layer. Output quality impact is still being tested.

Original post →

More from Infra

Infra channel →