Benchmarking Qwen 3.8-Flash-Next on Strix Halo: Halogen hits 1,045 prefill t/s

deepu105 · reddit · 2026-09-30

A systematic benchmark of inference engines running Qwen 3.8-Flash-Next on an AMD Strix Halo (128GB) at 70W TDP. Halogen 0.14.0 with native .hgn weights was fastest (1,045 prefill t/s, 39.3 decode t/s, 84% MTP accept), followed by open-source gufo (4x faster cold load, best time-to-first-token at 1.6-1.9s) and CIRU. Unsloth's llama.cpp builds lagged far behind at 285-314 prefill t/s. Methodology: 3-turn latency plus 1,000 output tokens, two runs each, with retrieval accuracy checks; top 3 engines were also tested at 128k/256k contexts.

Original post →

More from Infra

Infra channel →