Swapping Qwen's N-gram layer to Q8 shows no real speed penalty in local inference tests
Altruistic_Heat_9531 · reddit · 2026-09-03
A local-LLM experimenter replaced the low-precision N-gram portion of an IQ4XS Qwen 3.8 Next model with Q8 weights (following someone else's BF16 experiment) and benchmarked on a Xeon E5-2690v4 + RTX 3090 (250W cap) with 96GB DDR4, no MTP.
Short-run generation jumped from 8.8 to 10.7 tok/s, but both setups converge to a steady state around 10.1 tok/s — meaning no meaningful speed penalty from the higher-precision N-gram layer. Output quality impact is still being tested.
More from Infra
- Agentic API adds a stateful layer in front of vLLM for open-model agent runtimes — techNmak · 2026-09-03
- Google's Gemini 3.8 Flash 'works harder' but may burn more tokens at same pricing — The Verge AI · 2026-09-03
- Mitchell Hashimoto Details Memory Optimization Tricks in the Superlogical Server — sull · 2026-09-03
- Perplexity's Lily beats MLX-LM with 1.23x prefill and 1.35x decode throughput on M5 Max — perplexity_ai · 2026-09-03
- Perplexity open-sources Lily, a local inference engine for Qwen3.6 on Apple silicon — perplexity_ai · 2026-09-03
- FastH3 Now Runs Locally on Apple Silicon and DGX Spark — Vandy_simp · 2026-09-03