llama.cpp PR extends MoE fusion to speculative decoding, boosting MTP throughput
jacek2023 · reddit · 2026-08-31
A new llama.cpp PR (#27621 by ynankani) extends MoE GLU fusion and topk-router fusion—previously restricted to 1 token—to speculative decoding. The poster hasn't benchmarked it yet, but the included benchmarks suggest meaningful MTP speedups for MoE models across draft widths, especially greater than 1.
More from Infra
- TensorSharp vs llama.cpp: Qwen 3.8 Flash Next Benchmarks — fuzhongkai · 2026-09-01
- Why did increasing context size increase speed in Llama.cpp? — satnl · 2026-09-01
- AI inference demand surges again, supply brutally outpaced by token growth — Baconbrix · 2026-09-01
- Warp founder predicts cloud-based collaborative factories for all companies within a year — charlieholtz · 2026-09-01
- JPM: 1GW of AI Infrastructure Costs $40-45B, Frontier Labs Make ~$30B per GW — zephyr_z9 · 2026-09-01
- Data Center Worker: Fastest Blue-Collar Path to Six Figures Right Now — AICopyLab · 2026-09-01