Running Qwen Flash Next at 262k context on 96GB Strix Halo, seeking speedups
Forward_Jackfruit813 · reddit · 2026-09-05
The author runs Qwen Flash Next locally on a 96GB Strix Halo at IQ4XS with a 262k context limit, offloading NGRAMs to SSD. They're satisfied with the intelligence and now want more speed: 50PP/14Decode near full context. The post shares the full llama.cpp launch configuration (draft-mtp speculative decoding, f16 KV cache, flash-attn, unified KV) and asks whether an m.2-to-OCuLink V620 32GB eGPU would help on the Vulkan build, or if a 27B model might be a faster alternative.
More from Infra
- SGLang's Breakable CUDA Graph speeds prefill graph building by 3.8–5.2x — ying11231 · 2026-09-05
- SemiAnalysis: OpenAI's ASIC program is leverage — Altman wins even if the chip loses — MarvinTBaumann · 2026-09-05
- 2027 will be peak year of AI compute constraint; relief arrives in 2028, analyst argues — BenBajarin · 2026-09-05
- Dev builds P2P network to seed open-weight models, fearing a NVIDIA-Hugging Face deal — Relevant-Magic-Card · 2026-09-05
- NVIDIA lays out five practical guidelines for speculative decoding to speed up LLM inference — NVIDIAAI · 2026-09-05
- NVIDIA Details Hardware-Friendly LLM Design and Five Speculative Decoding Guidelines — NVIDIAAI · 2026-09-05