Lemonade fixes AMD APU model streaming, drops ROCm backend that was ~40x slower than Vulkan

Fcking_Chuck · reddit · 2026-09-24

Per Phoronix, AMD's Lemonade project (2026.40 RC) fixes model streaming on APUs. Notably, the release drops the OpenMOSS ROCm path, which benchmarked roughly 40x slower than the Vulkan backend on APUs — making Vulkan the de facto choice for local LLM inference on AMD APUs.

Original post →

More from Infra

Infra channel →