llama.cpp hits 1.2k t/s Qwen prefill on Strix Halo, matching closed-source Halogen
ilintar · reddit · 2026-09-13
A Reddit user pushed mainline llama.cpp's Qwen3.8 Flash Next prefill speed from 400 t/s to 1,200 t/s on AMD Strix Halo, matching closed-source server Halogen. They shipped a custom HIP runtime and install script, published an Opus-generated recap of the optimization journey, explained the mainline/fork ecosystem, and plan upstream PRs — work that should also benefit GLM 5.3 Flash's similar sparse attention.
More from Infra
- Yacine quips that Google's Astra is being load-shed — yacineMTB · 2026-09-13
- RAM crunch is real: 32GB VPS plans sold out as agent workloads pile on — flavioAd · 2026-09-13
- Hugging Face's $399 Microduck Sold 10,000 Units in Five Days, Powered by Shenzhen's Seeed — Scobleizer · 2026-09-13
- NVIDIA ships new AI factory architectures yearly at ~15x P/E; Apple's folding phone gets ~33x — JFPuget · 2026-09-13
- Full guide: run MiniMax H3 fast on AMD GPUs (RX 9070/R9700/7900) with ComfyUI and ROCm 7.14 — Apprehensive_Sky892 · 2026-09-13
- OpenAI data center to slash Georgia county property taxes by roughly 30% — surmenok · 2026-09-13