llama.cpp PR Uses AVX2 to Speed Up Large Batch IQ Quantization

pmttyji · reddit · 2026-08-20

A new PR in llama.cpp uses AVX2 instructions to significantly speed up inference for IQ-quantized models at large batch sizes. Benchmarks on Qwen3.6-27B and 35B-A3B show token processing speed improvements of 6x to over 12x for IQ series quantizations (e.g., IQ3S, IQ2XS) with negligible impact on perplexity.

Original post →

More from Infra

Infra channel →