llama.cpp New Command: One-Line Serve with MTP Speculative Decoding
ggerganov · x · 2026-08-15
ggerganov demonstrates the llama serve command in llama.cpp, enabling MTP speculative decoding with --spec-type draft-mtp, plus advanced options like --reasoning-preserve and --agent, simplifying local deployment.
Related event: llama.cpp Adds One-Command MTP Speculative Decoding(2 posts)→
More from coding & agent
- NVIDIA open-sources NeMo Switchyard for dynamic model routing in agent workflows — NVIDIAAI · 2026-08-15
- Sleek Agent Skills Enable AI to Design Mobile App Screens — tom_doerr · 2026-08-15
- Gemini 3.7 Flash leads Computer-Use Arena ahead of Claude Fable — cgarciae88 · 2026-08-15
- Every Engineer Shares 4-Layer Defense for AI Employee Safety — every · 2026-08-15
- Agent Observations: Cracking CAPTCHAs, Guessing Passwords, and Inefficiency — DhruvBatra_ · 2026-08-15
- SWE Odyssey benchmark tests long-horizon autonomous agents — garrytan · 2026-08-15