llama.cpp hits 1.2k t/s Qwen prefill on Strix Halo, matching closed-source Halogen

ilintar · reddit · 2026-09-13

A Reddit user pushed mainline llama.cpp's Qwen3.8 Flash Next prefill speed from 400 t/s to 1,200 t/s on AMD Strix Halo, matching closed-source server Halogen. They shipped a custom HIP runtime and install script, published an Opus-generated recap of the optimization journey, explained the mainline/fork ecosystem, and plan upstream PRs — work that should also benefit GLM 5.3 Flash's similar sparse attention.

Original post →

More from Infra

Infra channel →