MTP in llama.cpp now rivals ds4: GLM 5.3 Flash sets new Apple Silicon decode record
challis88ocarina · reddit · 2026-10-08
A Reddit user reports that MTP (Multi-Token Prediction) decoding in llama.cpp is now competitive with ds4 using GLM 5.3 Flash. Prompt processing is still slower, but this is the first time a model has outperformed ds4 in llama.cpp — even against a tuned M3U setup. MTP, which historically offered no advantage in llama.cpp, may finally be useful on Apple Silicon. The real test will be Qwen38FN, where vanilla ds4 currently manages 65 t/s (75 t/s concurrent).
More from Infra
- Tencent's STEPQuant: 6-bit recurrent states match FP32 with 68.7% less memory — _akhaliq · 2026-10-09
- Bain projects 38.6M GPU and custom silicon shipments by 2030, 10x 2023's 3.9M — Beth_Kindig · 2026-10-09
- AWS reference architecture: multi-team GPU cluster sharing on SageMaker HyperPod — AWS ML Blog · 2026-10-09
- Mistral slammed for training open models on datacenters powered ~70% by coal — wavefnx · 2026-10-09
- Why do we resend the whole conversation every turn? Server-side KV slots proposal sparks debate — Vasili_Sk · 2026-10-09
- NVIDIA's NeMo-DCR cuts 1T-model weight sync from 87.5 min to 150s, 12-40x faster checkpoint transfer — dair_ai · 2026-10-09