llama.cpp Metal gains M5 neural accelerator support, erasing MLX's prefill edge

MrPecunius · reddit · 2026-09-03

A Reddit user reports that llama.cpp's Metal backend appears to have gained support for the M5's matmul/neural accelerators, pushing GGUF prefill (e.g., Unsloth Q8) to 300-350 t/s on Apple M5 — on par with the best MLX results. Combined with GGUF+MTP already leading on token generation (19 t/s), the author sees no remaining reason to use MLX models, ending the need to switch formats based on prefill/generation mix. Tested with Qwen3 27B at Q8 on an M5 Pro; the post asks whether the community still sees a case for MLX or knows how to get mainstream MLX MTP working.

Original post →

More from Infra

Infra channel →