AMD Strix Halo user builds llama.cpp Laguna on ROCm and hits flash-attn limits

dbinnunE3 · reddit · 2026-07-22

A Reddit user documents a full build-and-benchmark attempt for the Laguna fork of llama.cpp on an AMD Strix Halo box using ROCm/HIP.

What they did

What worked

What failed

The post is essentially a detailed field report on getting a large local model stack running on AMD hardware and where the current rough edges are.

Original post →

More from coding & agent

coding & agent channel →