llama.cpp merges probabilistic MTP decoding, +14% speedup on prose generation
Dreeew84 · reddit · 2026-10-11
llama.cpp has merged a probabilistic MTP (multi-token prediction) decoding optimization (PR #27694), with reported 14% speedup on prose generation. The author recommends updating llama.cpp and trying it.
- Optimal draft-n-max / draft-p-min settings align with greedy sampling
- Main gains are on prose generation
- Tests ran with thinking and ngram-mod disabled
More from Infra
- Rent a GPU for ~€20/month: Gemma 4 31B hits 41 TPS on GeForce NOW via LM Studio — evilsocket · 2026-10-11
- Falling AI token prices are fueling demand — Jevons paradox at work, per a16z data — The Decoder · 2026-10-11
- MCIO cables extend PCIe at half-slot bandwidth, retimers fix signal loss — TheZachMueller · 2026-10-11
- Motherboard to switchboard won't populate your PCIe slots properly — TheZachMueller · 2026-10-11
- Cloudflare: sites blocked 13.47% of AI bot requests in Q3, more than double a year ago — YvesMulkers · 2026-10-11
- Final vllm-radiance Build Enables Multi-Agent on One R9700 GPU — KriptacMessage · 2026-10-11