177B Qwen 3.8 Flash Next Hits Steady ~45 tps Decode on a 128GB Strix Halo
deepu105 · reddit · 2026-10-08
Reddit user deepu105 reports that with the latest Halogen 0.17.2 release, running Qwen 3.8 Flash Next (177B parameters) locally on a 128GB Strix Halo now delivers a consistent 45 tps decode even at high context. The author praises peonist-ai's rapid commit pace, notes the model's quality on implementing and reviewing an "Opus 5.5 plan", and argues the community isn't appreciating this local-deployment capability enough.
More from Infra
- Abusers hop across inference providers, so providers must coordinate evictions — natolambert · 2026-10-08
- Enterprise AI trends toward model routers, not one giant model — ingliguori · 2026-10-08
- Crowdsourced harness x model benchmark: 3090 beats 5090 with qwen3.8-flash-next config — dh7net · 2026-10-08
- llama.cpp distributes inference across heterogeneous devices: MiMo 2.6 Flash at 40 tok/s over 10 GbE — joao_gante · 2026-10-08
- Java-based jitLLM claims 90% of llama.cpp perf on NVIDIA GPUs via TornadoVM CUDA compilation — mikebmx1 · 2026-10-08
- Taiwan OSATs Ordered $11B of Equipment in 9 Months, Topping 2019–2025 Combined — zephyr_z9 · 2026-10-08