MTP in llama.cpp now rivals ds4: GLM 5.3 Flash sets new Apple Silicon decode record

challis88ocarina · reddit · 2026-10-08

A Reddit user reports that MTP (Multi-Token Prediction) decoding in llama.cpp is now competitive with ds4 using GLM 5.3 Flash. Prompt processing is still slower, but this is the first time a model has outperformed ds4 in llama.cpp — even against a tuned M3U setup. MTP, which historically offered no advantage in llama.cpp, may finally be useful on Apple Silicon. The real test will be Qwen38FN, where vanilla ds4 currently manages 65 t/s (75 t/s concurrent).

Original post →

More from Infra

Infra channel →