llama.cpp Metal PR boosts speculative decoding up to 3.6x on Apple Silicon

pmttyji · reddit · 2026-10-06

A llama.cpp Metal pull request adds few-row MMA matmul and batched copies optimized for speculative decoding. On Apple Silicon with Qwen3.8-27B and a DFlash2 Q80 drafter, decode speed jumps from 30.2 to 110.0 tok/s on code and 16.8 to 62.6 tok/s on prose, with negligible overhead in serial mode. Notable for local LLM inference on Macs.

Original post →

More from Infra

Infra channel →