Qwen3.8-27B on Strix Halo: MTP gives 3.1x speedup, ROCm FP4 fork hits 29.9 tok/s

deepu105 · reddit · 2026-08-23

A rigorous local inference benchmark of Qwen3.8-27B on a Strix Halo laptop (Radeon 8060S, 128GB unified memory), using llama.cpp's own timing block, 3 reps per config, with 85W boost verified:

Headline results:

Key observations:

Runs driven via the author's LlamaStash tool with reproducible flags.

Original post →

More from Infra

Infra channel →