llama.cpp Metal gains M5 neural accelerator support, erasing MLX's prefill edge
MrPecunius · reddit · 2026-09-03
A Reddit user reports that llama.cpp's Metal backend appears to have gained support for the M5's matmul/neural accelerators, pushing GGUF prefill (e.g., Unsloth Q8) to 300-350 t/s on Apple M5 — on par with the best MLX results. Combined with GGUF+MTP already leading on token generation (19 t/s), the author sees no remaining reason to use MLX models, ending the need to switch formats based on prefill/generation mix. Tested with Qwen3 27B at Q8 on an M5 Pro; the post asks whether the community still sees a case for MLX or knows how to get mainstream MLX MTP working.
More from Infra
- Agentic API adds a stateful layer in front of vLLM for open-model agent runtimes — techNmak · 2026-09-03
- Mitchell Hashimoto Details Memory Optimization Tricks in the Superlogical Server — sull · 2026-09-03
- Perplexity's Lily beats MLX-LM with 1.23x prefill and 1.35x decode throughput on M5 Max — perplexity_ai · 2026-09-03
- Perplexity open-sources Lily, a local inference engine for Qwen3.6 on Apple silicon — perplexity_ai · 2026-09-03
- FastH3 Now Runs Locally on Apple Silicon and DGX Spark — Vandy_simp · 2026-09-03
- Claim: Without the datacenter buildout campaign, the US would be in a sharp recession — ZeeshanZiaML · 2026-09-03