llama.cpp PR claims 3-7x faster CPU prompt processing via VNNI

jacek2023 · reddit · 2026-09-26

An open llama.cpp pull request (#27851) by jbooth adds tiled mulmat for k-quants. Per the author, leveraging AVX512-VNNI delivers 3-7x faster CPU mulmat with minimal added complexity, substantially speeding up CPU-side prompt processing for local inference — a notable win for users without GPUs or limited by VRAM.

Original post →

More from Infra

Infra channel →