Running Qwen Flash Next at 262k context on 96GB Strix Halo, seeking speedups

Forward_Jackfruit813 · reddit · 2026-09-05

The author runs Qwen Flash Next locally on a 96GB Strix Halo at IQ4XS with a 262k context limit, offloading NGRAMs to SSD. They're satisfied with the intelligence and now want more speed: 50PP/14Decode near full context. The post shares the full llama.cpp launch configuration (draft-mtp speculative decoding, f16 KV cache, flash-attn, unified KV) and asks whether an m.2-to-OCuLink V620 32GB eGPU would help on the Vulkan build, or if a 27B model might be a faster alternative.

Original post →

More from Infra

Infra channel →