llama.cpp Adds One-Command MTP Speculative Decoding
llama.cpp introduced a new command that enables MTP speculative decoding with a single line, simplifying the setup for faster inference with models like Qwen3.
2026-08-15 ~ 2026-08-15 · 2 related posts
- llama.cpp new command: one-liner enables MTP speculative decoding — ggerganov · 2026-08-15
- llama.cpp New Command: One-Line Serve with MTP Speculative Decoding — ggerganov · 2026-08-15