llama.cpp PR extends MoE fusion to speculative decoding, boosting MTP throughput

jacek2023 · reddit · 2026-08-31

A new llama.cpp PR (#27621 by ynankani) extends MoE GLU fusion and topk-router fusion—previously restricted to 1 token—to speculative decoding. The poster hasn't benchmarked it yet, but the included benchmarks suggest meaningful MTP speedups for MoE models across draft widths, especially greater than 1.

Original post →

More from Infra

Infra channel →