llama.cpp PR Enables CUDA Graphs for MTP Draft, Delivering Another Inference Speedup

jacek2023 · reddit · 2026-09-17

A new llama.cpp pull request (#28549 by a NVIDIA engineer) enables CUDA Graphs for the MTP draft path, reducing kernel launch overhead and adding another speedup to speculative decoding with multi-token prediction.

Original post →

More from Infra

Infra channel →