llama.cpp lands Flash Attention tuning for RDNA4, big prefill gains on AMD

pmttyji · reddit · 2026-09-11

llama.cpp merged pwilkin's CUDA/HIP Flash Attention tuning PR (#28102) targeting AMD RDNA4 (R9700) and RDNA 3.5 (RX 9060 XT, 8060S) GPUs, delivering notable prefill improvements with especially good numbers at large context. Detailed benchmarks are included in the PR — a real speedup for running LLMs locally on AMD cards.

Original post →

More from Infra

Infra channel →