llama.cpp Adds Q8_0 Support
pmttyji · reddit · 2026-07-15
ggml-zendnn added Q80 quantization support to llama.cpp and compared the throughput performance of GGML CPU versus ZenDNN across multiple models.
Key Takeaways
- During the prompt processing phase, ZenDNN shows significant improvements, especially with longer prompts.
- The decode phase (tg128) is basically on par with ggml-cpu.
- Tested models include Llama-3.1-8B-Instruct, Mixtral-8x7B, gemma4 31B, and gemma-4-26B-A4B-it.
Representative Data
- Llama-3.1-8B-Instruct: Increased from 472.28 t/s to 730.87 t/s at 256 prompt length, a 54.75% boost; the increase approaches 92% at 2048 prompt length.
- Mixtral-8x7B: Increased from 150.11 t/s to 470.41 t/s at 2048 prompt length, a 213.38% boost.
- gemma4 31B: Increased from 106.37 t/s to 222.32 t/s at 2048 prompt length, a 109.01% boost.
- gemma-4-26B-A4B-it: Relatively smaller improvement, around 14.20% at 2048 prompt length.
Conclusion
- ZenDNN significantly accelerates the preprocessing phase for long prompts.
- Such optimizations are biased towards inference stacks and local deployment performance, rather than changing the model's inherent capabilities.
More from Infra
- RTK Terminal Compression Cuts Tokens but Leaves Your AI Coding Bill Unchanged — Bartaseth · 2026-09-11
- SF Compute founder: buying compute is 'an absolutely awful experience' right now — IgorCarron · 2026-09-11
- SmolVM open-sources persistent computer infrastructure for agents that outlive chat sessions — aniketmaurya · 2026-09-11
- PyTorch Day Korea 2026 launches first offline conf, CFP closes Sept 13 — PyTorch · 2026-09-11
- Local LLM server dilemma: 4x CMP-170HX (price up 53% in 20 days) vs Mac Studio M5 Ultra — rumboll · 2026-09-11
- llama.cpp lands Flash Attention tuning for RDNA4, big prefill gains on AMD — pmttyji · 2026-09-11