llama.cpp Adds Q8_0 Support
pmttyji · reddit · 2026-07-15
ggml-zendnn added Q80 quantization support to llama.cpp and compared the throughput performance of GGML CPU versus ZenDNN across multiple models.
Key Takeaways
- During the prompt processing phase, ZenDNN shows significant improvements, especially with longer prompts.
- The decode phase (tg128) is basically on par with ggml-cpu.
- Tested models include Llama-3.1-8B-Instruct, Mixtral-8x7B, gemma4 31B, and gemma-4-26B-A4B-it.
Representative Data
- Llama-3.1-8B-Instruct: Increased from 472.28 t/s to 730.87 t/s at 256 prompt length, a 54.75% boost; the increase approaches 92% at 2048 prompt length.
- Mixtral-8x7B: Increased from 150.11 t/s to 470.41 t/s at 2048 prompt length, a 213.38% boost.
- gemma4 31B: Increased from 106.37 t/s to 222.32 t/s at 2048 prompt length, a 109.01% boost.
- gemma-4-26B-A4B-it: Relatively smaller improvement, around 14.20% at 2048 prompt length.
Conclusion
- ZenDNN significantly accelerates the preprocessing phase for long prompts.
- Such optimizations are biased towards inference stacks and local deployment performance, rather than changing the model's inherent capabilities.
More from Infra
- Agent search bottlenecks are now about variance, not raw latency — rohanpaul_ai · 2026-07-22
- Gavin Baker argues Nvidia may be one of open source AI’s biggest supporters — GavinSBaker · 2026-07-22
- AI Power Demand Exposes US Energy Gap, Urging Shift from Scarcity to Abundance — bradneuberg · 2026-07-22
- Gavin Baker says Nvidia’s $630B figure would be system revenue, not all Nvidia’s — GavinSBaker · 2026-07-22
- A Firecracker-based platform says it can host 6,000 AI agents on one 256 GB server — maritime_sh · 2026-07-22
- Report says Nvidia could build 1,000 Vera Rubin racks a day, implying $630B quarterly at system level — GavinSBaker · 2026-07-22