AMD Users Report: Official llama.cpp Branch Boosts PP Speed Over 2x

PromptInjection_ · reddit · 2026-08-23

A user highlights AMD's officially maintained branch of llama.cpp, noting it is actively maintained despite deprecation warnings (changes are later upstreamed). Testing on a Strix Halo with the ROCm/Hip backend revealed that Prompt Processing speed for dense models more than doubled (rising from 230 to 550 t/s for a 14B dense model). However, Text Generation (TG) speed is about 15% slower than with Vulkan. MoE model speeds remain unchanged.

Original post →

More from Infra

Infra channel →