llama.cpp PR Uses AVX2 to Speed Up Large Batch IQ Quantization
pmttyji · reddit · 2026-08-20
A new PR in llama.cpp uses AVX2 instructions to significantly speed up inference for IQ-quantized models at large batch sizes. Benchmarks on Qwen3.6-27B and 35B-A3B show token processing speed improvements of 6x to over 12x for IQ series quantizations (e.g., IQ3S, IQ2XS) with negligible impact on perplexity.
More from Infra
- Opposition to local data centers in US surges 33 points to 75% — Polymarket · 2026-08-21
- Ramp Router cuts GPT-5.6 Sol inference costs by 50% — KlausCodes · 2026-08-21
- NVIDIA releases official CUDA MCP server for AI-assisted dev — swagonflyyyy · 2026-08-21
- Escha claims 2bit quant matches FP8 performance in benchmarks — luedtek · 2026-08-21
- Researcher Rants: Conference Season Blocks GPU Access for Days — ChongZzZhang · 2026-08-21
- AT&T routes 40% of employee AI usage to open models — Hesamation · 2026-08-21