llama.cpp new command: one-liner enables MTP speculative decoding
ggerganov · x · 2026-08-15
Georgi Gerganov, author of llama.cpp, demonstrated a simple command on X: llama serve -hf ggml-org/Qwen3.8-27B-GGUF --spec-type draft-mtp to enable MTP (Multi-Token Prediction) speculative decoding. This simplifies the workflow for using speculative decoding, which is practical for optimizing local inference performance.
Related event: llama.cpp Adds One-Command MTP Speculative Decoding(2 posts)→
More from coding & agent
- GitHub Copilot adds Grok 4.6, Kimi K3, and other new models — film_girl · 2026-08-15
- Developer confirms hiding data in code comments works to fool models — suchenzang · 2026-08-15
- Comic 4 render speedup test: significantly faster than legacy tools — andrew_n_carr · 2026-08-15
- WebMCP demo: Agents can call website tools directly via URL — jasonkneen · 2026-08-15
- Hermes Agents Integrates @skills Hub for Direct Skill Addition — Teknium · 2026-08-15
- Tutorial: Building MCP servers for databases — adnan_hashmi · 2026-08-15