llama.cpp Adds MTP Support for Qwen3-Next, Enabling Full-Speed Inference

jacek2023 · reddit · 2026-08-03

A newly merged PR #25589 in llama.cpp introduces MTP (Multi-Token Prediction) support for the Qwen3-Next model.

This means developers and users can now run Qwen3-Next locally at "full speed," significantly boosting generation efficiency.

Original post →

More from Infra

Infra channel →