llama.cpp merges probabilistic MTP decoding, +14% speedup on prose generation

Dreeew84 · reddit · 2026-10-11

llama.cpp has merged a probabilistic MTP (multi-token prediction) decoding optimization (PR #27694), with reported 14% speedup on prose generation. The author recommends updating llama.cpp and trying it.

Original post →

More from Infra

Infra channel →