Radix-select top-k fix boosts Qwen Flash-Next decode 9-12% at ~119k context on 2x3090
Extension-Bid-639 · reddit · 2026-09-08
An update on running Qwen3 Flash-Next with expert cache and MTP on dual RTX 3090s: replacing the CUDA top-k fallback (which sorted whole rows due to missing newer CUB DeviceTopK) with llama.cpp's existing radix-selection code — matching upstream PR #28366 — cut top-k kernel time from 5.1ms to 0.25ms per committed token.
Setup: 2x RTX 3090, dual Xeon E5-2696 v4, 128GB DDR4, UD-Q4KXL quants, f16 KV, 150 expert-cache slots, MTP-3, 261,888 allocated context, long-document tests starting at 119k tokens. Median decode went from 30.2-30.4 t/s to 33.1-33.7 t/s across three seeds — 9-12% higher per-seed median throughput.
A quality screen of 480 requests (240 matched comparisons across 80 question/depth combos and 3 seeds) showed control 235/240 correct vs candidate 238/240; the author notes this only disproves regression, not prove improvement. The change is now in the author's production build; prefill and no-MTP gains remain unmeasured. Users on older CUDA toolkits running Flash-Next should check their top-k fallback.
More from Infra
- FT: CXMT and YMTC stockpiled enough ASML DUV tools for three years of expansion — basedjensen · 2026-09-08
- Nvidia's $59.7B quarterly profit tops 12 iconic companies combined ($57.9B) — FinanceYF5 · 2026-09-08
- Consolidating four small models into one inference server: a doc-QA agent's ops tradeoffs — Sad-Razzmatazz-7657 · 2026-09-08
- Huawei Ascend takes PyTorch China stage: from hardware adaptation to joint standards research — PyTorch · 2026-09-08
- AUR llama.cpp-cuda removal caused 10x slowdown; manual rebuild restored 1800 t/s prefill — MrHall · 2026-09-08
- CPO laser series part 3: MOPA approach vs Lumentum's single-cavity route — vikramskr · 2026-09-08