llama.cpp Adds Q2_0 Quantization Support

pmttyji · reddit · 2026-07-08

This PR introduces CPU-side Q20 quantization support to ggml/llama.cpp, targeting Ternary Bonsai 1.7B, 4B, 8B, and other 1.58-bit models, as well as future ones. The implementation currently covers CPU only, including ARM NEON and generic scalar fallback.

Original post →

More from Infra

Infra channel →