ROCm Beats Vulkan on Strix Halo: Up to 47% Faster Decoding at 41.9 tok/s

QuixiAI · x · 2026-09-17

The Lucebox team reports pushing ROCm past Vulkan on AMD Strix Halo. Benchmarked same-day on the same machine with identical prompts: with the DSpark drafter, Lucebox (ROCm) decodes at 41.9 tok/s at 8K and 38.0 tok/s at 123K prompt tokens — 28-47% faster than llama.cpp/Vulkan v0.7.5, with 17-47% faster prefill. A notable win for local inference on AMD's AI Max hardware.

Original post →

More from Infra

Infra channel →