llama.cpp PR adds MTP support for Qwen Flash Next, GGUF quants now on Hugging Face
jacek2023 · reddit · 2026-10-01
A merged llama.cpp PR (#29761) adds MTP (multi-token prediction) support for Qwen Flash Next via the Qwen4Exp branch, improving local inference speed. GGUF quants are available at ggml-org/Qwen3.8-Flash-Next-GGUF on Hugging Face.
More from Infra
- Case study: how Canva saved millions in cloud costs — _jaydeepkarale · 2026-10-01
- tilelang: A DSL for High-Performance GPU/Accelerator Kernels Hits 7,973 Stars — tile-ai · 2026-10-01
- Bain: AI companies need $4.2 trillion a year in new revenue by 2031 to fund data centers — ylecun · 2026-10-01
- Dev ditches AWS vector service for open-source Weaviate after scaling pain — CShorten30 · 2026-10-01
- Nebius acquires Inferize to cut the GPU idle tax in production inference — demian_ai · 2026-10-01
- One chart of the AI infra buildout: 2026 construction boom, 2030 bet on demand — AntDX316 · 2026-10-01