llama.cpp lands three Metal MoE PRs, decode jumps from 65.6 to 73.9 tok/s

predatar · reddit · 2026-09-03

An M5 MacBook Pro owner obsessed with not cooking their laptop submitted 3 Metal backend PRs to llama.cpp and is asking Reddit for testers, especially on Qwen3.8-Flash-Next:

Benchmark via llama-bench -m model.gguf -p 4096 -n 64 -ub 512 comparing master vs the PR.

Original post →

More from Infra

Infra channel →