MTP released for Qwen3.8-Flash-Next GGUF, promising big local TPS gains

vini542reddit · reddit · 2026-09-01

MTP (Multi-Token Prediction) support has been released for the GGUF quantization of Qwen3.8-Flash-Next. The poster expects it to significantly boost local tokens-per-second and is eager to test it.

What remains is for more llama.cpp optimizations to be merged so local users can benefit.

Original post →

More from Infra

Infra channel →