AMD llama.cpp patch boosts context length from 64K to 149K for Qwen 27B

ea_man · reddit · 2026-08-09

Developer eaman shares a patch for llama.cpp that reduces MTP buffer overhead, significantly increasing available context length. On ROCm and Vulkan backends, Qwen 27B context length jumps from 64K to 149K (dual GPU). Patch and logs are public.

Original post →

More from Infra

Infra channel →