MTP released for Qwen3.8-Flash-Next GGUF, promising big local TPS gains
vini542reddit · reddit · 2026-09-01
MTP (Multi-Token Prediction) support has been released for the GGUF quantization of Qwen3.8-Flash-Next. The poster expects it to significantly boost local tokens-per-second and is eager to test it.
What remains is for more llama.cpp optimizations to be merged so local users can benefit.
More from Infra
- Ollama Switches to Transparent Per-Token Pricing with Monthly Credit Pools — ollama · 2026-09-01
- Omarchy achieves first Linux TouchID crack on T1 MacBooks — DanWahlin · 2026-09-01
- Question: Have data providers started training their own models? — xeophon · 2026-09-01
- Stop leaving your AI Agent running 24/7: Power management guide for developers — Rhishi99 · 2026-09-01
- Huge price gaps in Token resources: self-deployed GLM and DeepSeek available at up to 80% off — lipeng0820 · 2026-09-01
- Whale's May paper introduced multi-plane network architecture, influencing AI training and chip design — bookwormengr · 2026-09-01