llama.cpp PR Boosts Q2_0 CPU Decoding by 3x, 8B Hits 8.20 tok/s

BTA_Labs · reddit · 2026-08-07

A new llama.cpp PR (#26348) introduces an x86 VNNI implementation for the Q20 × Q80 dot product, achieving a massive 3.0x to 3.6x throughput improvement on x86 CPUs.

Original post →

More from Infra

Infra channel →