Qwen3.8-Flash-Next (125B) runs at 59 tok/s on a single Strix Halo mini PC, engine open-sourced
Yaniss916 · reddit · 2026-10-05
Yamz got Qwen3.8-Flash-Next (125B MoE, 6B active) running well on a single AMD Strix Halo mini PC (Ryzen AI Max+ 395, 128GB), releasing 95GB EXL3 weights and the open Kyojin inference engine built on ExLlamaV3.
- Speed: 44–59 tok/s with speculative decoding (47 chat, 58 code); 32.7 tok/s without. A llama.cpp user reported 30 tok/s on the same box.
- Prefill: 1,412 tok/s at 4K, nearly flat across contexts; 10/10 needle tests at 64K and 128K, still 32 tok/s at 128K.
- Fidelity: 94.1% top-1 agreement with the original FP8 model vs Halogen's 92.3%, and 41% lower KL divergence — though Halogen 0.16.2 is faster overall; the team is reworking the engine core to close the gap.
- Speculative decoding is verified to return exactly the tokens plain decoding would; an optional uncensor preset ships off by default.
More from Infra
- NVIDIA Dynamo lets coding agents point at self-hosted endpoints with native tracing — TheZachMueller · 2026-10-06
- Schmidhuber: compute gets 10x cheaper every 5 years, 100,000x in 25 years — SchmidhuberAI · 2026-10-06
- Musk confirms TSMC in talks to build chips at his planned Texas Terafab — Polymarket · 2026-10-06
- Running MiniMax H3 locally on a 12GB GPU: 0.6MP at 4:3 is the sweet spot — vortis23 · 2026-10-06
- AI Data Center Boom Hits Grid Limits as Planned Projects Get Pulled Back — DavidLinthicum · 2026-10-05
- Kirin 9050 reportedly uses 1.5-micron HBI die-to-wafer packaging; analyst says scaling it is the hard part — teortaxesTex · 2026-10-05