AMD Users Report: Official llama.cpp Branch Boosts PP Speed Over 2x
PromptInjection_ · reddit · 2026-08-23
A user highlights AMD's officially maintained branch of llama.cpp, noting it is actively maintained despite deprecation warnings (changes are later upstreamed). Testing on a Strix Halo with the ROCm/Hip backend revealed that Prompt Processing speed for dense models more than doubled (rising from 230 to 550 t/s for a 14B dense model). However, Text Generation (TG) speed is about 15% slower than with Vulkan. MoE model speeds remain unchanged.
More from Infra
- Mistral reportedly plans up to 1 GW of European compute capacity by 2030 — emmanuelvivier · 2026-08-23
- Contextual News Search APIs: A Deep Comparison for AI, RAG, and Research — ermanos12 · 2026-08-23
- Qwen3.8-27B MTP Grafted to Unsloth Saves RAM, Requires Thinking Mode — Nyghtbynger · 2026-08-23
- Running Kimi K3 on 8x B300: $190 per million tokens, full cost breakdown — OtherRaisin3426 · 2026-08-23
- Nvidia AI Server Prices to Rise 15% Due to DRAM Shortage — The Decoder · 2026-08-23
- ComfyUI Node Optimization: Sparse Attention Boosts Speed by 5-20% — Zironic · 2026-08-23