Llama.cpp on ROCm Crashes Constantly on Radeon 780M—One Env Var Fixes It

MaximusSenior · reddit · 2026-08-27

Running llama.cpp with ROCm 7.14 on a Radeon 780M iGPU with Qwen 3, the author found small-prompt preprocessing hits 200-300 t/s (vs 60 t/s on Vulkan), but with frequent crashes.

After tracing it to a known upstream ROCm issue, the workaround is setting AMDSERIALIZEKERNEL=3: much more stable, at the cost of preprocessing dropping to 100 t/s—still faster than Vulkan. The author asks whether others on the 780M or similar iGPUs hit the same issue or found a better fix.

Original post →

More from Infra

Infra channel →