gfx906-llama-cpp Fork Boosts MI50/MI60 Inference +23% Prefill, Fits 250k Context on 40GB

milpster · reddit · 2026-09-05

Developer milpster shipped a major update to gfx906-llama-cpp, a fork optimizing llama.cpp for AMD GCN cards like MI50/MI60/Radeon VII: prefill PP16384 up from 332.5 to 410 t/s (+23%), 120k deep fill +14%, TG decode +11% (15.1 t/s), and 250k context squeezed into 40GB VRAM (upstream can't fit it). Outputs are bit-identical to upstream. Gains came mostly from adopting relevant upstream PRs; the author notes the fork was built with AI assistance. A practical resource for cheap local LLM inference on aging GPUs.

Original post →

More from Infra

Infra channel →