Qwen3.8-Flash-Next (125B) runs at 59 tok/s on a single Strix Halo mini PC, engine open-sourced

Yaniss916 · reddit · 2026-10-05

Yamz got Qwen3.8-Flash-Next (125B MoE, 6B active) running well on a single AMD Strix Halo mini PC (Ryzen AI Max+ 395, 128GB), releasing 95GB EXL3 weights and the open Kyojin inference engine built on ExLlamaV3.

Original post →

More from Infra

Infra channel →