llama.cpp Metal PR boosts speculative decoding up to 3.6x on Apple Silicon
pmttyji · reddit · 2026-10-06
A llama.cpp Metal pull request adds few-row MMA matmul and batched copies optimized for speculative decoding. On Apple Silicon with Qwen3.8-27B and a DFlash2 Q80 drafter, decode speed jumps from 30.2 to 110.0 tok/s on code and 16.8 to 62.6 tok/s on prose, with negligible overhead in serial mode. Notable for local LLM inference on Macs.
More from Infra
- Nvidia has a moat but not control: Google's TPUs now hold a quarter of global AI compute — AryHHAry · 2026-10-06
- Local LLM reality check: 4090 caps Qwen context at ~32k before VRAM runs out — BLUECOW009 · 2026-10-06
- DDR5 Overclocking Boosts Out-of-VRAM MoE Inference by Up to 14% in Real Tests — EvolvingDior · 2026-10-06
- vLLM's Transformers backend now serves text, image, audio and video with zero custom code — LysandreJik · 2026-10-06
- DeepSeek raises at least $12B, eyes IPO in early 2027, Bloomberg reports — teortaxesTex · 2026-10-06
- SEC filing: 80% of Anthropic's $518B compute commitments owed even if idle — rohanpaul_ai · 2026-10-06