Qwen Flash Next MTP work resumes with official GGUF quants and llama.cpp PR
jacek2023 · reddit · 2026-10-01
MTP (multi-token prediction) work for Qwen Flash Next has restarted. llama.cpp users can switch to the official ggml-org Qwen3.8-Flash-Next-GGUF quantized files on Hugging Face, with support landing via llama.cpp PR #29761. The author notes it is still work in progress, but the MTP acceleration path for local deployments is being restored.
More from Infra
- Arduino argues sub-$900 embedded boards beat Mac minis for Physical AI agent economics — CatAstro_Piyush · 2026-10-01
- SemiAnalysis injects failures to rate GPU cluster renters — providers differ sharply on recovery — AccBalanced · 2026-10-01
- Redditor runs unattended DeepSeek loops for days: 237M tokens for just $3.48 — dogfoodarchitect · 2026-10-01
- Free course built from Cornell's GPU architecture workshop now shared publicly — idanbeck · 2026-10-01
- Undocumented Strata tip: set default sampling params via a sampling block in run config — KissMyShinyArse · 2026-10-01
- Meta's Loop Scaling Laws: Sparsity Gives ~3x Active-Param Efficiency, Recurrence ~2x on Reasoning — facebook · 2026-10-01