llama.cpp PR claims 3-7x faster CPU prompt processing via VNNI
jacek2023 · reddit · 2026-09-26
An open llama.cpp pull request (#27851) by jbooth adds tiled mulmat for k-quants. Per the author, leveraging AVX512-VNNI delivers 3-7x faster CPU mulmat with minimal added complexity, substantially speeding up CPU-side prompt processing for local inference — a notable win for users without GPUs or limited by VRAM.
More from Infra
- NVIDIA at $5.4 trillion is now worth more than the entire UK or French stock market — iamfakhrealam · 2026-09-26
- DeepSeek V4.1 Flash's Engram memory layer trades FFN compute for lookup tables, SemiAnalysis data suggests — teortaxesTex · 2026-09-26
- A Curated Paper List for Learning Distributed LLM Training and Inference — East-Muffin-6472 · 2026-09-26
- vLLM adds Elastic Expert Parallelism: grow/shrink MoE GPU pools under live traffic — PyTorch · 2026-09-26
- AI data center investors now favor real infrastructure over PowerPoint promises — TansuYegen · 2026-09-26
- Running Qwen 27B Q4_K_M on dual RTX 3060 with llama.cpp hits ~44-50 tok/s — jacek2023 · 2026-09-26