llama.cpp Metal optimization boosts IQ3_XXS decode speed on Apple Silicon

predatar · reddit · 2026-09-01

A submitted PR optimizes the Metal backend in llama.cpp, improving decode speed for IQ3XXS models on Apple Silicon. Benchmarks show TieL Coder 35B A3B decode speed increased from 65.6 to 73.9 tok/s. The author invites community testing and teases upcoming prefill improvements.

Original post →

More from Infra

Infra channel →