AMD llama.cpp patch boosts context length from 64K to 149K for Qwen 27B
ea_man · reddit · 2026-08-09
Developer eaman shares a patch for llama.cpp that reduces MTP buffer overhead, significantly increasing available context length. On ROCm and Vulkan backends, Qwen 27B context length jumps from 64K to 149K (dual GPU). Patch and logs are public.
More from Infra
- LeCun Weighs In: Breakthrough AI Silicon Fails Without Software Ecosystem — ylecun · 2026-08-09
- Running LLMs on Snapdragon NPUs: A Guide to Qualcomm's GenieX CLI — carrycooldude · 2026-08-09
- AI Compute Costs: UK Datacentre Expansion Sparks Water and Power Crises — nordicinst · 2026-08-09
- Developer Creates MiniMax H3 RunPod Template for Easy Deployment — Draufgaenger · 2026-08-09
- Amazon's Planned Gas Plant for AI Data Center Could Become Top US Climate Polluter — Nunki08 · 2026-08-09
- UltraEP: Near-Optimal Load Balancing for Rack-Scale MoE Training — jiqizhixin · 2026-08-09