llama.cpp new command: one-liner enables MTP speculative decoding

ggerganov · x · 2026-08-15

Georgi Gerganov, author of llama.cpp, demonstrated a simple command on X: llama serve -hf ggml-org/Qwen3.8-27B-GGUF --spec-type draft-mtp to enable MTP (Multi-Token Prediction) speculative decoding. This simplifies the workflow for using speculative decoding, which is practical for optimizing local inference performance.

Related event: llama.cpp Adds One-Command MTP Speculative Decoding(2 posts)→

Original post →

More from coding & agent

coding & agent channel →