llama.cpp merges new Metal kernels, making speculative decoding 3.4x faster on M3 Ultra
ggerganov · x · 2026-10-05
- A PR from the qvac team, reviewed and merged personally by llama.cpp author ggerganov, adds new Metal kernels to fix speculative decoding being slower than plain decoding on Apple Silicon.
- Problem: speculative decoding and batched decoding produce mat-muls with only 2–16 rows; Apple GPUs (M1–M4) lack the tensor API, so Metal handled these with mat-vec kernels whose time grows per row — on an M3 Ultra, DFlash2 decoding of Qwen3 27B was slower than serial decoding on master.
- Fix: new mat-mul kernels for 2–16 src1 rows built on 8x8 simdgroup matrices; weights are dequantized once for all rows, with simdgroups in a threadgroup splitting K. Dedicated kernels for Q40, Q80 and Q5K.
- Result: on an M3 Ultra, speculative decoding now hits 110 tok/s vs 32.1 tok/s for plain decoding — roughly 3.4x faster.
More from coding & agent
- 7 Claude Prompts to Learn Anything Faster, From Planning to Flashcards — CodeByPoonam · 2026-10-05
- Matt Pocock Ships AI Coding Skills v1.3: /implement-spec, /pr, /retro and GLOSSARY.md — mattpocockuk · 2026-10-05
- Pi adds MCP support after publicly mocking it—engineering team explains what changed — bibryam · 2026-10-05
- "Don't stop until perfect" instruction burned $450 in Azure egress fees on a cloud agent — pswider · 2026-10-05
- Open-source personal agent Comma argues PAs need sessionless continuity across tasks — Tricky_Barnacle_2060 · 2026-10-05
- Agent-Reach hits 91K GitHub stars giving AI agents cookie-based, zero-fee web access with ban risks — mhdfaran · 2026-10-05