Llama.cpp on ROCm Crashes Constantly on Radeon 780M—One Env Var Fixes It
MaximusSenior · reddit · 2026-08-27
Running llama.cpp with ROCm 7.14 on a Radeon 780M iGPU with Qwen 3, the author found small-prompt preprocessing hits 200-300 t/s (vs 60 t/s on Vulkan), but with frequent crashes.
After tracing it to a known upstream ROCm issue, the workaround is setting AMDSERIALIZEKERNEL=3: much more stable, at the cost of preprocessing dropping to 100 t/s—still faster than Vulkan. The author asks whether others on the 780M or similar iGPUs hit the same issue or found a better fix.
More from Infra
- Pushing for LoRA sharing to reduce download waste — Borkato · 2026-08-27
- Local Deployment of GLM-5.3-Flash: 206 tok/s and 1M Context on DGX Station — funding__secured · 2026-08-27
- Hugging Face launches Jobs: run UV/Docker workloads on any hardware, pay per second — _akhaliq · 2026-08-27
- Merge, PostHog, and Redis host NYC technical talks on self-driving AI products — shensi · 2026-08-27
- LightningAI offers instant H100 access on its self-owned AI cloud — LightningAI · 2026-08-27
- M7 Ultra may feature native FP8, potentially boosting GLM 5.3-flash performance — Brilliant-Hall1387 · 2026-08-27