177B Qwen 3.8 Flash Next Hits Steady ~45 tps Decode on a 128GB Strix Halo

deepu105 · reddit · 2026-10-08

Reddit user deepu105 reports that with the latest Halogen 0.17.2 release, running Qwen 3.8 Flash Next (177B parameters) locally on a 128GB Strix Halo now delivers a consistent 45 tps decode even at high context. The author praises peonist-ai's rapid commit pace, notes the model's quality on implementing and reviewing an "Opus 5.5 plan", and argues the community isn't appreciating this local-deployment capability enough.

Original post →

More from Infra

Infra channel →