llama.cpp lands three Metal MoE PRs, decode jumps from 65.6 to 73.9 tok/s
predatar · reddit · 2026-09-03
An M5 MacBook Pro owner obsessed with not cooking their laptop submitted 3 Metal backend PRs to llama.cpp and is asking Reddit for testers, especially on Qwen3.8-Flash-Next:
- PR #28301: general MoE prefill optimization that skips empty work in underfilled expert tiles, plus an IQ2/IQ3-specific dequant path that matters most on M5 (IQ3XXS)
- PR #28302: fixes checkpoint eviction for hybrid/recurrent models so editing, branching or reopening sessions resumes from the recent KV/recurrent state instead of re-prefilling a large chunk
- PR #28086: IQ3XXS Metal decode optimization — on Tiel-Coder-35B-A3B, decode went from 65.6 to 73.9 tok/s
Benchmark via llama-bench -m model.gguf -p 4096 -n 64 -ub 512 comparing master vs the PR.
More from Infra
- Nvidia acquires Hugging Face, the 'GitHub of AI,' for $13 billion — Hakan_Ozalp · 2026-09-03
- Figure commits $3.5B (scaling past $6B) to deploy up to 100,000 GPUs with Nscale — Scobleizer · 2026-09-03
- Hidden China risks in America's multibillion-dollar AI data center boom are well known, insider says — pstAsiatech · 2026-09-03
- Nvidia nears $12.9B Hugging Face acquisition, about 86x revenue — sanjaykalra · 2026-09-03
- Reka and NVIDIA unveil real-time 30B video model: 720p at 24fps, 11.8x faster on one H100 — RekaAILabs · 2026-09-03
- AI Data Center 'Powered Shell' Stocks: Q2 Earnings Recap and What Could Reignite Interest — BenBajarin · 2026-09-03