llama.cpp PR Enables CUDA Graphs for MTP Draft, Delivering Another Inference Speedup
jacek2023 · reddit · 2026-09-17
A new llama.cpp pull request (#28549 by a NVIDIA engineer) enables CUDA Graphs for the MTP draft path, reducing kernel launch overhead and adding another speedup to speculative decoding with multi-token prediction.
More from Infra
- How Bell Labs Missed the Microchip: IEEE Spectrum Revisits a Landmark Tech-History Blunder — ArtificialOther · 2026-09-17
- Agentic AI systems are the next network users: 40% of enterprise apps to include agents by 2026 — seankinneyRCR · 2026-09-17
- B200 spot rental up 80% in 8 months as demand outpaces compute buildout — JOBhakdi · 2026-09-17
- MLX-Serve 26.9.3 ships: Qwen Flash Next tops 100 tok/s on M4/M5 Max Macs — TheMoonMidas · 2026-09-17
- Running a 14B Model on 16GB RAM: 'My PC Is a Toaster Now' — Aggravating_Site381 · 2026-09-17
- fal engineering head: we'll never pre-train, inference compute is the real moat — jfischoff · 2026-09-17