Benchmarking Qwen 3.8-Flash-Next on Strix Halo: Halogen hits 1,045 prefill t/s
deepu105 · reddit · 2026-09-30
A systematic benchmark of inference engines running Qwen 3.8-Flash-Next on an AMD Strix Halo (128GB) at 70W TDP. Halogen 0.14.0 with native .hgn weights was fastest (1,045 prefill t/s, 39.3 decode t/s, 84% MTP accept), followed by open-source gufo (4x faster cold load, best time-to-first-token at 1.6-1.9s) and CIRU. Unsloth's llama.cpp builds lagged far behind at 285-314 prefill t/s. Methodology: 3-turn latency plus 1,000 output tokens, two runs each, with retrieval accuracy checks; top 3 engines were also tested at 128k/256k contexts.
More from Infra
- Claude+small-model harness cuts 1,000 AI decisions from $605 to $0.17 — jamestagg · 2026-09-30
- Intel hires new data center sales VP from Marvell — firstadopter · 2026-09-30
- PSA: Power-limit your RTX 6000 Pro to the minimum — inference perf unaffected — AlpinDale · 2026-09-30
- Analyst: Neo Clouds Aren't Winning Enterprise Customers, They're Trading GPUs Among Themselves — DavidLinthicum · 2026-09-30
- a16z's State of Markets: hyperscaler capex nears $1 trillion a year — a16z Podcast · 2026-09-30
- Cloudflare's MCP redesign cuts tool context from 244K to 1.1K via catalog + executor — PuzzledFarmer4554 · 2026-09-30