halogen 0.12.0 hits 38 tok/s decode at 1M context for Qwen on Strix Halo
peonist-ai · reddit · 2026-09-20
peonist-ai released halogen-flash-server 0.12.0, fixing context-depth performance degradation for Qwen3.8-Flash-Next on AMD's Ryzen AI Max+ 395 (128GB Strix Halo).
- Decode at 1,004,581 tokens of context improved from 27.3 to 38.3 tok/s (default speculative drafter); 42.9→45.0 tok/s at 258,794 context
- 1M prefill rose from 790 to 937 tok/s, cold-request time cut from 21.2 to 17.9 minutes; 256k prefill now 1,114 tok/s
- Follow-up turns over the prompt cache hit first token in 0.55s at 1M context
- Setup: add HALOGENROPEYARN=4 and HALOGENCTX=1048576 to the podman line; needs a 128GB box; open source on GitHub
More from Infra
- Qwen3.8 Flash Next on one RTX 5090 hits 50 t/s decode via FreeToken expert caching — dir3ctly · 2026-09-20
- Are data centers dodging taxes? Tax Foundation data on $1B facilities says no — AndyMasley · 2026-09-20
- AMD carries out serious software optimizations for Kimi-K3, analyst says — AccBalanced · 2026-09-20
- RAM price jumps from $350 to $640 in months, pricing out new PCs — chrisalbon · 2026-09-20
- NVIDIA engineer breaks down why DeepSeek re-engineered V4.1 Flash for speed — thursdai_pod · 2026-09-20
- $140 Radeon MI50 paired with GTX-1080Ti boosts local 27B-35B LLM speeds up to 9x — tabletuser_blogspot · 2026-09-20