llama.cpp PR adds MTP support for Qwen Flash Next, GGUF quants now on Hugging Face

jacek2023 · reddit · 2026-10-01

A merged llama.cpp PR (#29761) adds MTP (multi-token prediction) support for Qwen Flash Next via the Qwen4Exp branch, improving local inference speed. GGUF quants are available at ggml-org/Qwen3.8-Flash-Next-GGUF on Hugging Face.

Original post →

More from Infra

Infra channel →