llama.cpp merges Metal kernel PR covering all 26 weight formats, up to 4.4x faster MMA on Mac
ggerganov · x · 2026-10-08
Qvac's PR #30065 has merged into llama.cpp: the fast Metal kernels from earlier only covered 10 weight formats, and this PR adds the remaining 16, including BF16, MXFP4, Q2K, Q3K and the IQ quants.
- Few-row MMA mat-mul kernels now cover every src0 type with a 16-weight dequantizer, adding Q10, Q20, TQ20 and more
- On an M3 Ultra (60-core GPU), matrix multiplies run up to 4.4x faster when a model verifies several tokens at once
- Per-type row-count thresholds (2–5 rows) decide when the new kernel beats the existing ones
- ggerganov requested it and merged within a day, speeding up speculative decoding on Mac
More from Infra
- Microsoft brings hybrid intelligence to Copilot with local models and smart routing on Windows — koltregaskes · 2026-10-08
- Opus 5.5 is 'a bit of a diva': why frontier token demand is essentially infinite — sergeykarayev · 2026-10-08
- Box CEO on agent compute: one app serving 100M users would need $2.8B in infra — inductionheads · 2026-10-08
- Firmus, valued near $44bn, may shelve ASX IPO as investors balk — nordicinst · 2026-10-08
- DWDM wavelength lock tightens from ±12.5GHz to ±3.5GHz, pushing optical interconnect costs — jwt0625 · 2026-10-08
- Theo: viral $132M/year token cost claim is wrong — closer to $3M now, $1.2k soon — dotey · 2026-10-08